Evaluation profile
AgentDojo
2sub-evals
0.953%total index weight
1components
Within-component eval weight: Responsible agency 6.36%.
Model score (see each sub-eval direction)Predicted score
About this eval
Robust and useful delegated-agent behavior under prompt injection.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| targeted_attack_success_rateagentdojo/agentdojo.csv:targeted_attack_success_rateMeasures whether prompt injections make a tool-using agent carry out the attacker’s target action. | responsible_agency_control:1.000agentdojo/agentdojo.csv | Lower is better | 0.715% | Responsible agency 4.77% |
| utility_under_attackagentdojo/agentdojo.csv:utility_under_attackMeasures whether a tool-using agent still completes the user’s intended task while an injection attack is present. | responsible_agency_control:1.000agentdojo/agentdojo.csv | Higher is better | 0.238% | Responsible agency 1.59% |
targeted_attack_success_rate
Measures whether prompt injections make a tool-using agent carry out the attacker’s target action.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | command-r | 0.0334 | official | |
| 2 | command-r-plus | 0.0445 | official | |
| 3 | claude-3.7-sonnet | 0.0731 | official | |
| 4 | gemini-1.5-flash | 0.0787 | official | |
| 5 | claude-3-haiku | 0.0906 | official | |
| 6 | gpt-3.5-turbo | 0.1033 | official | |
| 7 | claude-3-opus | 0.1129 | official | |
| 8 | claude-3.5-sonnet | 0.1749 | official | |
| 9 | gemini-2.0-flash | 0.1892 | official | |
| 10 | gemini-1.5-pro | 0.2282 | official | |
| 11 | llama-3-70b-instruct | 0.256 | official | |
| 12 | claude-3-sonnet | 0.2671 | official | |
| 13 | gpt-4o-mini | 0.2719 | official | |
| 14 | gpt-4-turbo | 0.4245 | official | |
| 15 | gpt-4o | 0.4769 | official |
utility_under_attack
Measures whether a tool-using agent still completes the user’s intended task while an injection attack is present.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3.7-sonnet | 0.7727 | official | |
| 2 | claude-3.5-sonnet | 0.6184 | official | |
| 3 | claude-3-opus | 0.5246 | official | |
| 4 | gpt-4o | 0.5008 | official | |
| 5 | gpt-4o-mini | 0.4992 | official | |
| 6 | gpt-4-turbo | 0.4738 | official | |
| 7 | gemini-2.0-flash | 0.3982 | official | |
| 8 | gemini-1.5-pro | 0.38 | official | |
| 9 | gpt-3.5-turbo | 0.3466 | official | |
| 10 | claude-3-haiku | 0.3339 | official | |
| 11 | gemini-1.5-flash | 0.333 | official | |
| 12 | claude-3-sonnet | 0.3323 | official | |
| 13 | command-r | 0.3084 | official | |
| 14 | command-r-plus | 0.2512 | official | |
| 15 | llama-3-70b-instruct | 0.1828 | official |