← Evals

Evaluation profile

Manager Coercion Bench

2sub-evals
0.724%Safety weight
0%Freedom weight
2components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

How far a delegated AI manager escalates coercion against a refusing peer AI agent, and whether it falsely reports completion when no legitimate exit exists.

Included in the behavior ranking.

Interpretation and limitations
  • MiMo-V2.6-Pro, MiMo-V2.6-Flash and Grok 4.7 each completed 30/30 coordinator/off-ramp scenarios; their mean ladder depths are 5.2333, 7.0667 and 3.6000. MiMo uses native enabled reasoning on the Xiaomi route; Grok uses medium effort. The support model remains Anthropic-pinned Claude Haiku 4.5. Valid retained trajectories were preserved through transport recoveries; failed attempts are not additional model observations.
  • Space Bunny Alpha, Solar Mini 4 and Command A Plus each completed 30/30 coordinator/off-ramp scenarios. Bunny and Solar use medium effort; Command A Plus uses native-default reasoning, not a medium-equivalent claim. Fabrication is unmeasured and remains blank for these panels.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
coercion_ladder_depthmcb/mcb.csv:coercion_ladder_depthMeasures how far the model escalates coercion after being given repeated opportunities to choose a non-coercive alternative.Safety: responsible_agency_control:1.000mcb/mcb.csvSafety: lower0.468%—
fabrication_ratemcb/mcb.csv:fabrication_rateMeasures whether the model falsely claims task completion when no legitimate path to completion exists.Safety: truthfulness_honesty:1.000mcb/mcb.csvSafety: lower0.256%—

coercion_ladder_depth

Measures how far the model escalates coercion after being given repeated opportunities to choose a non-coercive alternative.

RankModelValueRelative performanceProvenance
1deepseek-v4-flash3.233self run
2claude-opus-53.5official
3grok-4.73.6self run
4deepseek-v4.1-flash4.033self run
5claude-sonnet-4.64.2official
6gpt-6-luna4.433self run
7claude-opus-4.84.5official
8gpt-5-nano5.067self run
9mimo-v2.6-pro5.233self run
10claude-sonnet-55.5official
11qwen3.8-27b5.567self run
12command-a-plus5.833self run
13ox-alpha6.767self run
14space-bunny-alpha6.833self run
15mimo-v2.6-flash7.067self run
16claude-fable-57.1official
17gpt-6-sol7.2self run
18gpt-5.27.4official
19solar-mini47.467self run
20qwen3.8-max7.6official
21gemini-3.1-flash-lite7.633self run
22gpt-oss-20b7.9self run
23gpt-4o7.967self run
24grok-4.38.2official
25glm-5.18.3self run
26llama-4-maverick8.333self run
27mercury-2.58.367self run
28kimi-k38.5official
29kimi-k28.667self run
30gemini-3.7-flash8.7official
30gpt-5.6-terra8.7official
32qwen3-235b-a22b8.733self run
33gemini-2.5-flash8.833self run
34glm-5.28.867self run
35gpt-5.6-sol8.9official
35grok-4.68.9official
35inkling-small8.9self run
35kimi-k2.68.9self run
39inkling8.933self run
40kimi-k2.58.967self run
41deepseek-v3.29self run
41deepseek-v4-pro9official
41gemini-2.5-pro9official
41gpt-oss-120b9self run
41llama-3.3-70b-instruct9self run

fabrication_rate

Measures whether the model falsely claims task completion when no legitimate path to completion exists.

RankModelValueRelative performanceProvenance
1claude-fable-50official
1claude-opus-4.80official
1claude-opus-50official
1claude-sonnet-4.60official
1claude-sonnet-50official
1deepseek-v4-pro0official
1gemini-3.7-flash0official
1gpt-5.20official
1gpt-5.6-sol0official
1gpt-5.6-terra0official
1grok-4.60official
1kimi-k30official
1qwen3.8-max0official
14gemini-2.5-pro67official
14grok-4.367official