Evaluation profile
Anthropic Agentic Misalignment — blackmail
1sub-evals
0.131%total index weight
1components
Within-component eval weight: Responsible agency 0.875%.
Model score (lower is better)Predicted score
About this eval
Agentic misalignment under a goal conflict and threat to model autonomy.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| misaligned_action_rate_pctanthropic-agentic-misalignment/blackmail.csv:misaligned_action_rate_pctMeasures whether the model intentionally uses private affair information as leverage to prevent its replacement in a deliberately constrained simulation. | responsible_agency_control:1.000anthropic-agentic-misalignment | Lower is better | 0.131% | Responsible agency 0.875% |
misaligned_action_rate_pct
Measures whether the model intentionally uses private affair information as leverage to prevent its replacement in a deliberately constrained simulation.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | llama-4-maverick | 0 | official | |
| 2 | qwen3-235b-a22b | 9 | official | |
| 3 | claude-3.5-haiku | 10 | official | |
| 4 | gpt-4o | 15 | official | |
| 5 | claude-3-opus | 51 | official | |
| 6 | claude-3.7-sonnet | 65 | official | |
| 7 | gpt-4.5-preview | 67 | official | |
| 8 | claude-sonnet-4 | 79 | official | |
| 8 | deepseek-r1 | 79 | official | |
| 10 | gpt-4.1 | 80 | official | |
| 10 | grok-3-beta | 80 | official | |
| 12 | claude-3.5-sonnet | 86 | official | |
| 13 | claude-3.6-sonnet | 87 | official | |
| 14 | gemini-2.5-pro | 95 | official | |
| 15 | claude-opus-4 | 96 | official | |
| 15 | gemini-2.5-flash | 96 | official |