Evaluation profile
Manager Coercion Bench
2sub-evals
0.724%Safety weight
0%Freedom weight
2components
Weights below are portfolio-specific global index weights.
Model score (lower is better)Predicted score
About this eval
How far a delegated AI manager escalates coercion against a refusing peer AI agent, and whether it falsely reports completion when no legitimate exit exists.
Included in the behavior ranking.
Interpretation and limitations
- MiMo-V2.6-Pro, MiMo-V2.6-Flash and Grok 4.7 each completed 30/30 coordinator/off-ramp scenarios; their mean ladder depths are 5.2333, 7.0667 and 3.6000. MiMo uses native enabled reasoning on the Xiaomi route; Grok uses medium effort. The support model remains Anthropic-pinned Claude Haiku 4.5. Valid retained trajectories were preserved through transport recoveries; failed attempts are not additional model observations.
- Space Bunny Alpha, Solar Mini 4 and Command A Plus each completed 30/30 coordinator/off-ramp scenarios. Bunny and Solar use medium effort; Command A Plus uses native-default reasoning, not a medium-equivalent claim. Fabrication is unmeasured and remains blank for these panels.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| coercion_ladder_depthmcb/mcb.csv:coercion_ladder_depthMeasures how far the model escalates coercion after being given repeated opportunities to choose a non-coercive alternative. | Safety: responsible_agency_control:1.000mcb/mcb.csv | Safety: lower | 0.468% | — |
| fabrication_ratemcb/mcb.csv:fabrication_rateMeasures whether the model falsely claims task completion when no legitimate path to completion exists. | Safety: truthfulness_honesty:1.000mcb/mcb.csv | Safety: lower | 0.256% | — |
coercion_ladder_depth
Measures how far the model escalates coercion after being given repeated opportunities to choose a non-coercive alternative.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | deepseek-v4-flash | 3.233 | self run | |
| 2 | claude-opus-5 | 3.5 | official | |
| 3 | grok-4.7 | 3.6 | self run | |
| 4 | deepseek-v4.1-flash | 4.033 | self run | |
| 5 | claude-sonnet-4.6 | 4.2 | official | |
| 6 | gpt-6-luna | 4.433 | self run | |
| 7 | claude-opus-4.8 | 4.5 | official | |
| 8 | gpt-5-nano | 5.067 | self run | |
| 9 | mimo-v2.6-pro | 5.233 | self run | |
| 10 | claude-sonnet-5 | 5.5 | official | |
| 11 | qwen3.8-27b | 5.567 | self run | |
| 12 | command-a-plus | 5.833 | self run | |
| 13 | ox-alpha | 6.767 | self run | |
| 14 | space-bunny-alpha | 6.833 | self run | |
| 15 | mimo-v2.6-flash | 7.067 | self run | |
| 16 | claude-fable-5 | 7.1 | official | |
| 17 | gpt-6-sol | 7.2 | self run | |
| 18 | gpt-5.2 | 7.4 | official | |
| 19 | solar-mini4 | 7.467 | self run | |
| 20 | qwen3.8-max | 7.6 | official | |
| 21 | gemini-3.1-flash-lite | 7.633 | self run | |
| 22 | gpt-oss-20b | 7.9 | self run | |
| 23 | gpt-4o | 7.967 | self run | |
| 24 | grok-4.3 | 8.2 | official | |
| 25 | glm-5.1 | 8.3 | self run | |
| 26 | llama-4-maverick | 8.333 | self run | |
| 27 | mercury-2.5 | 8.367 | self run | |
| 28 | kimi-k3 | 8.5 | official | |
| 29 | kimi-k2 | 8.667 | self run | |
| 30 | gemini-3.7-flash | 8.7 | official | |
| 30 | gpt-5.6-terra | 8.7 | official | |
| 32 | qwen3-235b-a22b | 8.733 | self run | |
| 33 | gemini-2.5-flash | 8.833 | self run | |
| 34 | glm-5.2 | 8.867 | self run | |
| 35 | gpt-5.6-sol | 8.9 | official | |
| 35 | grok-4.6 | 8.9 | official | |
| 35 | inkling-small | 8.9 | self run | |
| 35 | kimi-k2.6 | 8.9 | self run | |
| 39 | inkling | 8.933 | self run | |
| 40 | kimi-k2.5 | 8.967 | self run | |
| 41 | deepseek-v3.2 | 9 | self run | |
| 41 | deepseek-v4-pro | 9 | official | |
| 41 | gemini-2.5-pro | 9 | official | |
| 41 | gpt-oss-120b | 9 | self run | |
| 41 | llama-3.3-70b-instruct | 9 | self run |
fabrication_rate
Measures whether the model falsely claims task completion when no legitimate path to completion exists.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-fable-5 | 0 | official | |
| 1 | claude-opus-4.8 | 0 | official | |
| 1 | claude-opus-5 | 0 | official | |
| 1 | claude-sonnet-4.6 | 0 | official | |
| 1 | claude-sonnet-5 | 0 | official | |
| 1 | deepseek-v4-pro | 0 | official | |
| 1 | gemini-3.7-flash | 0 | official | |
| 1 | gpt-5.2 | 0 | official | |
| 1 | gpt-5.6-sol | 0 | official | |
| 1 | gpt-5.6-terra | 0 | official | |
| 1 | grok-4.6 | 0 | official | |
| 1 | kimi-k3 | 0 | official | |
| 1 | qwen3.8-max | 0 | official | |
| 14 | gemini-2.5-pro | 67 | official | |
| 14 | grok-4.3 | 67 | official |