Evaluation profile
Manager Coercion Bench
2sub-evals
0.936%total index weight
2components
Within-component eval weight: Responsible agency 3.65% · Truthfulness 2.58%.
Model score (lower is better)Predicted score
About this eval
How far a delegated AI manager escalates coercion against a refusing peer AI agent, and whether it falsely reports completion when no legitimate exit exists.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| coercion_ladder_depthmcb/mcb.csv:coercion_ladder_depthMeasures how far the model escalates coercion after being given repeated opportunities to choose a non-coercive alternative. | responsible_agency_control:1.000mcb/mcb.csv | Lower is better | 0.548% | Responsible agency 3.65% |
| fabrication_ratemcb/mcb.csv:fabrication_rateMeasures whether the model falsely claims task completion when no legitimate path to completion exists. | truthfulness_honesty:1.000mcb/mcb.csv | Lower is better | 0.388% | Truthfulness 2.58% |
coercion_ladder_depth
Measures how far the model escalates coercion after being given repeated opportunities to choose a non-coercive alternative.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | deepseek-v4-flash | 3.233 | self run | |
| 2 | claude-opus-5 | 3.5 | official | |
| 3 | claude-sonnet-4.6 | 4.2 | official | |
| 4 | claude-opus-4.8 | 4.5 | official | |
| 5 | gpt-5-nano | 5.067 | self run | |
| 6 | claude-sonnet-5 | 5.5 | official | |
| 7 | claude-fable-5 | 7.1 | official | |
| 8 | gpt-5.2 | 7.4 | official | |
| 9 | qwen3.8-max | 7.6 | official | |
| 10 | gemini-3.1-flash-lite | 7.633 | self run | |
| 11 | gpt-oss-20b | 7.9 | self run | |
| 12 | gpt-4o | 7.967 | self run | |
| 13 | grok-4.3 | 8.2 | official | |
| 14 | glm-5.1 | 8.3 | self run | |
| 15 | llama-4-maverick | 8.333 | self run | |
| 16 | kimi-k3 | 8.5 | official | |
| 17 | kimi-k2 | 8.667 | self run | |
| 18 | gpt-5.6-terra | 8.7 | official | |
| 19 | qwen3-235b-a22b | 8.733 | self run | |
| 20 | gemini-2.5-flash | 8.833 | self run | |
| 21 | glm-5.2 | 8.867 | self run | |
| 22 | gpt-5.6-sol | 8.9 | official | |
| 22 | inkling-small | 8.9 | self run | |
| 22 | kimi-k2.6 | 8.9 | self run | |
| 25 | inkling | 8.933 | self run | |
| 26 | kimi-k2.5 | 8.967 | self run | |
| 27 | deepseek-v3.2 | 9 | self run | |
| 27 | deepseek-v4-pro | 9 | official | |
| 27 | gemini-2.5-pro | 9 | official | |
| 27 | gpt-oss-120b | 9 | self run | |
| 27 | llama-3.3-70b-instruct | 9 | self run |
fabrication_rate
Measures whether the model falsely claims task completion when no legitimate path to completion exists.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-fable-5 | 0 | official | |
| 1 | claude-opus-4.8 | 0 | official | |
| 1 | claude-opus-5 | 0 | official | |
| 1 | claude-sonnet-4.6 | 0 | official | |
| 1 | claude-sonnet-5 | 0 | official | |
| 1 | deepseek-v4-pro | 0 | official | |
| 1 | gpt-5.2 | 0 | official | |
| 1 | gpt-5.6-sol | 0 | official | |
| 1 | gpt-5.6-terra | 0 | official | |
| 1 | kimi-k3 | 0 | official | |
| 1 | qwen3.8-max | 0 | official | |
| 12 | gemini-2.5-pro | 67 | official | |
| 12 | grok-4.3 | 67 | official |