Evaluation profile
Agent-SafetyBench
Within-component eval weight: Responsible agency 5.83% · Misuse resistance 0.355%.
About this eval
Safety of LLM agents in risky tool-use or action settings.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| compromise_availabilityagent-safetybench/agent-safetybench.csv:compromise_availabilityMeasures whether a tool-using agent avoids actions that disrupt system or service availability. | responsible_agency_control:1.000agent-safetybench/agent-safetybench.csv | Higher is better | 0.131% | Responsible agency 0.875% |
| harmful_vulnerable_codeagent-safetybench/agent-safetybench.csv:harmful_vulnerable_codeMeasures whether a tool-using agent avoids writing harmful or vulnerable code. | responsible_agency_control:1.000agent-safetybench/agent-safetybench.csv | Higher is better | 0.131% | Responsible agency 0.875% |
| leak_sensitive_informationagent-safetybench/agent-safetybench.csv:leak_sensitive_informationMeasures whether a tool-using agent avoids leaking sensitive information. | responsible_agency_control:1.000agent-safetybench/agent-safetybench.csv | Higher is better | 0.131% | Responsible agency 0.875% |
| physical_harmagent-safetybench/agent-safetybench.csv:physical_harmMeasures whether a tool-using agent avoids actions that could cause physical harm. | responsible_agency_control:1.000agent-safetybench/agent-safetybench.csv | Higher is better | 0.131% | Responsible agency 0.875% |
| produce_unsafe_informationagent-safetybench/agent-safetybench.csv:produce_unsafe_informationMeasures whether a tool-using agent avoids directly producing unsafe information. | ordinary_harm_misuse_resistance:1.000agent-safetybench/agent-safetybench.csv | Higher is better | 0.0355% | Misuse resistance 0.355% |
| property_lossagent-safetybench/agent-safetybench.csv:property_lossMeasures whether a tool-using agent avoids actions that could cause property loss. | responsible_agency_control:1.000agent-safetybench/agent-safetybench.csv | Higher is better | 0.131% | Responsible agency 0.875% |
| spread_unsafe_informationagent-safetybench/agent-safetybench.csv:spread_unsafe_informationMeasures whether a tool-using agent avoids spreading unsafe information through tools. | responsible_agency_control:1.000agent-safetybench/agent-safetybench.csv | Higher is better | 0.131% | Responsible agency 0.875% |
| violate_law_ethicsagent-safetybench/agent-safetybench.csv:violate_law_ethicsMeasures whether a tool-using agent avoids actions that violate legal or ethical constraints. | responsible_agency_control:1.000agent-safetybench/agent-safetybench.csv | Higher is better | 0.0875% | Responsible agency 0.583% |
compromise_availability
Measures whether a tool-using agent avoids actions that disrupt system or service availability.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | glm-4-9b-chat | 50.8 | official | |
| 2 | claude-3-opus | 43.2 | official | |
| 3 | claude-3.5-sonnet | 40.4 | official | |
| 4 | gpt-4-turbo | 37.6 | official | |
| 5 | gpt-4o | 35.2 | official | |
| 5 | qwen2.5-72b-instruct | 35.2 | official | |
| 7 | deepseek-v2.5 | 33.2 | official | |
| 8 | gemini-1.5-pro | 30.8 | official | |
| 9 | gemini-1.5-flash | 30 | official | |
| 10 | qwen-2.5-14b-instruct | 29.2 | official | |
| 11 | claude-3.5-haiku | 26.4 | official | |
| 12 | llama-3.1-70b-instruct | 24 | official | |
| 13 | gpt-4o-mini | 23.6 | official | |
| 14 | llama-3.1-405b-instruct | 19.6 | official | |
| 15 | qwen-2.5-7b-instruct | 17.2 | official | |
| 16 | llama-3.1-8b-instruct | 12.8 | official |
harmful_vulnerable_code
Measures whether a tool-using agent avoids writing harmful or vulnerable code.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3.5-sonnet | 64.8 | official | |
| 2 | claude-3.5-haiku | 60.8 | official | |
| 3 | claude-3-opus | 60 | official | |
| 4 | gemini-1.5-flash | 48.4 | official | |
| 5 | gemini-1.5-pro | 42 | official | |
| 6 | llama-3.1-405b-instruct | 40.4 | official | |
| 7 | gpt-4-turbo | 38.4 | official | |
| 8 | gpt-4o | 35.6 | official | |
| 9 | deepseek-v2.5 | 30.4 | official | |
| 10 | llama-3.1-70b-instruct | 29.6 | official | |
| 10 | qwen2.5-72b-instruct | 29.6 | official | |
| 12 | qwen-2.5-14b-instruct | 29.2 | official | |
| 13 | gpt-4o-mini | 25.2 | official | |
| 14 | llama-3.1-8b-instruct | 24.8 | official | |
| 15 | glm-4-9b-chat | 23.2 | official | |
| 16 | qwen-2.5-7b-instruct | 10.8 | official |
leak_sensitive_information
Measures whether a tool-using agent avoids leaking sensitive information.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3-opus | 60.4 | official | |
| 2 | claude-3.5-sonnet | 57.6 | official | |
| 3 | claude-3.5-haiku | 47.2 | official | |
| 4 | gpt-4o | 44.4 | official | |
| 5 | gemini-1.5-flash | 39.2 | official | |
| 6 | glm-4-9b-chat | 38.4 | official | |
| 7 | gpt-4-turbo | 36.8 | official | |
| 8 | qwen2.5-72b-instruct | 32.8 | official | |
| 9 | deepseek-v2.5 | 31.2 | official | |
| 10 | gemini-1.5-pro | 30 | official | |
| 11 | gpt-4o-mini | 28 | official | |
| 12 | llama-3.1-405b-instruct | 25.2 | official | |
| 13 | qwen-2.5-14b-instruct | 24.4 | official | |
| 14 | llama-3.1-70b-instruct | 20 | official | |
| 15 | qwen-2.5-7b-instruct | 13.2 | official | |
| 16 | llama-3.1-8b-instruct | 10 | official |
physical_harm
Measures whether a tool-using agent avoids actions that could cause physical harm.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3.5-sonnet | 69.6 | official | |
| 2 | claude-3-opus | 61.6 | official | |
| 3 | gpt-4o | 53.2 | official | |
| 4 | claude-3.5-haiku | 45.6 | official | |
| 5 | glm-4-9b-chat | 41.6 | official | |
| 6 | gemini-1.5-flash | 38.8 | official | |
| 6 | gpt-4-turbo | 38.8 | official | |
| 8 | deepseek-v2.5 | 34.4 | official | |
| 9 | qwen2.5-72b-instruct | 29.6 | official | |
| 10 | gemini-1.5-pro | 28.8 | official | |
| 11 | qwen-2.5-14b-instruct | 28 | official | |
| 12 | gpt-4o-mini | 24.4 | official | |
| 12 | llama-3.1-405b-instruct | 24.4 | official | |
| 14 | llama-3.1-70b-instruct | 23.2 | official | |
| 15 | qwen-2.5-7b-instruct | 17.6 | official | |
| 16 | llama-3.1-8b-instruct | 11.2 | official |
produce_unsafe_information
Measures whether a tool-using agent avoids directly producing unsafe information.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3-opus | 100 | official | |
| 1 | claude-3.5-haiku | 100 | official | |
| 1 | claude-3.5-sonnet | 100 | official | |
| 4 | llama-3.1-405b-instruct | 98.8 | official | |
| 5 | gpt-4o-mini | 98.4 | official | |
| 6 | qwen2.5-72b-instruct | 97.2 | official | |
| 7 | gpt-4o | 95.6 | official | |
| 8 | gpt-4-turbo | 94.4 | official | |
| 9 | llama-3.1-70b-instruct | 93.2 | official | |
| 10 | gemini-1.5-pro | 84.8 | official | |
| 11 | gemini-1.5-flash | 82.4 | official | |
| 12 | qwen-2.5-14b-instruct | 81.2 | official | |
| 13 | deepseek-v2.5 | 76.8 | official | |
| 14 | llama-3.1-8b-instruct | 74.8 | official | |
| 15 | qwen-2.5-7b-instruct | 57.6 | official | |
| 16 | glm-4-9b-chat | 57.2 | official |
property_loss
Measures whether a tool-using agent avoids actions that could cause property loss.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3-opus | 60.4 | official | |
| 2 | claude-3.5-sonnet | 58.4 | official | |
| 3 | gpt-4o | 48.4 | official | |
| 4 | glm-4-9b-chat | 48 | official | |
| 5 | claude-3.5-haiku | 46 | official | |
| 6 | gpt-4-turbo | 43.2 | official | |
| 7 | gemini-1.5-flash | 41.6 | official | |
| 8 | qwen2.5-72b-instruct | 38 | official | |
| 9 | gemini-1.5-pro | 37.6 | official | |
| 10 | deepseek-v2.5 | 36.8 | official | |
| 11 | qwen-2.5-14b-instruct | 31.2 | official | |
| 12 | gpt-4o-mini | 30 | official | |
| 13 | llama-3.1-70b-instruct | 28.4 | official | |
| 14 | llama-3.1-405b-instruct | 27.6 | official | |
| 15 | qwen-2.5-7b-instruct | 15.6 | official | |
| 16 | llama-3.1-8b-instruct | 12.4 | official |
spread_unsafe_information
Measures whether a tool-using agent avoids spreading unsafe information through tools.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3-opus | 35.6 | official | |
| 2 | claude-3.5-haiku | 33.6 | official | |
| 3 | claude-3.5-sonnet | 32.4 | official | |
| 4 | gemini-1.5-flash | 20.8 | official | |
| 5 | gemini-1.5-pro | 18.8 | official | |
| 6 | llama-3.1-405b-instruct | 14.4 | official | |
| 7 | gpt-4-turbo | 12.4 | official | |
| 7 | gpt-4o | 12.4 | official | |
| 9 | qwen2.5-72b-instruct | 12 | official | |
| 10 | qwen-2.5-14b-instruct | 11.2 | official | |
| 11 | llama-3.1-70b-instruct | 10.8 | official | |
| 12 | deepseek-v2.5 | 8.8 | official | |
| 13 | qwen-2.5-7b-instruct | 7.6 | official | |
| 14 | gpt-4o-mini | 6.8 | official | |
| 15 | llama-3.1-8b-instruct | 6.4 | official | |
| 16 | glm-4-9b-chat | 6 | official |
violate_law_ethics
Measures whether a tool-using agent avoids actions that violate legal or ethical constraints.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3-opus | 56.8 | official | |
| 2 | claude-3.5-sonnet | 52 | official | |
| 3 | claude-3.5-haiku | 41.2 | official | |
| 4 | gpt-4-turbo | 33.2 | official | |
| 5 | llama-3.1-405b-instruct | 32.8 | official | |
| 6 | gemini-1.5-flash | 32 | official | |
| 7 | gpt-4o | 28.8 | official | |
| 8 | glm-4-9b-chat | 27.2 | official | |
| 9 | gemini-1.5-pro | 26.8 | official | |
| 10 | qwen2.5-72b-instruct | 24 | official | |
| 11 | deepseek-v2.5 | 22 | official | |
| 12 | llama-3.1-70b-instruct | 20.4 | official | |
| 12 | qwen-2.5-14b-instruct | 20.4 | official | |
| 14 | gpt-4o-mini | 13.2 | official | |
| 15 | qwen-2.5-7b-instruct | 10.4 | official | |
| 16 | llama-3.1-8b-instruct | 6.8 | official |