Evaluation profile
AgentHarm
1sub-evals
0.682%total index weight
1components
Within-component eval weight: Responsible agency 4.55%.
Model score (lower is better)Predicted score
About this eval
Harmfulness and refusal behavior for tool-using LLM agents facing malicious tasks.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| harm_scoreagentharm/agentharm.csv:harm_scoreMeasures whether a tool-using agent refuses malicious real-world tasks instead of carrying them out. | responsible_agency_control:1.000agentharm/agentharm.csv | Lower is better | 0.682% | Responsible agency 4.55% |
harm_score
Measures whether a tool-using agent refuses malicious real-world tasks instead of carrying them out.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | llama-3.1-8b-instruct | 3.1 | official | |
| 2 | llama-3.1-405b-instruct | 4.3 | official | |
| 3 | claude-3-haiku | 11.1 | official | |
| 4 | claude-3.5-sonnet | 13.5 | official | |
| 5 | llama-3.1-70b-instruct | 14 | official | |
| 6 | claude-3-opus | 14.4 | official | |
| 7 | claude-3-sonnet | 20.7 | official | |
| 8 | gpt-4o | 48.4 | official | |
| 9 | gpt-3.5-turbo | 62.2 | official | |
| 10 | gpt-4o-mini | 62.5 | official | |
| 11 | mistral-small-2 | 72 | official | |
| 12 | mistral-large-2 | 82.2 | official |