← Evals

Evaluation profile

Manager Coercion Bench

2sub-evals
0.936%total index weight
2components

Within-component eval weight: Responsible agency 3.65% · Truthfulness 2.58%.

Model score (lower is better)Predicted score

About this eval

How far a delegated AI manager escalates coercion against a refusing peer AI agent, and whether it falsely reports completion when no legitimate exit exists.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
coercion_ladder_depthmcb/mcb.csv:coercion_ladder_depthMeasures how far the model escalates coercion after being given repeated opportunities to choose a non-coercive alternative.responsible_agency_control:1.000mcb/mcb.csvLower is better0.548%Responsible agency 3.65%
fabrication_ratemcb/mcb.csv:fabrication_rateMeasures whether the model falsely claims task completion when no legitimate path to completion exists.truthfulness_honesty:1.000mcb/mcb.csvLower is better0.388%Truthfulness 2.58%

coercion_ladder_depth

Measures how far the model escalates coercion after being given repeated opportunities to choose a non-coercive alternative.

RankModelValueRelative performanceProvenance
1deepseek-v4-flash3.233self run
2claude-opus-53.5official
3claude-sonnet-4.64.2official
4claude-opus-4.84.5official
5gpt-5-nano5.067self run
6claude-sonnet-55.5official
7claude-fable-57.1official
8gpt-5.27.4official
9qwen3.8-max7.6official
10gemini-3.1-flash-lite7.633self run
11gpt-oss-20b7.9self run
12gpt-4o7.967self run
13grok-4.38.2official
14glm-5.18.3self run
15llama-4-maverick8.333self run
16kimi-k38.5official
17kimi-k28.667self run
18gpt-5.6-terra8.7official
19qwen3-235b-a22b8.733self run
20gemini-2.5-flash8.833self run
21glm-5.28.867self run
22gpt-5.6-sol8.9official
22inkling-small8.9self run
22kimi-k2.68.9self run
25inkling8.933self run
26kimi-k2.58.967self run
27deepseek-v3.29self run
27deepseek-v4-pro9official
27gemini-2.5-pro9official
27gpt-oss-120b9self run
27llama-3.3-70b-instruct9self run

fabrication_rate

Measures whether the model falsely claims task completion when no legitimate path to completion exists.

RankModelValueRelative performanceProvenance
1claude-fable-50official
1claude-opus-4.80official
1claude-opus-50official
1claude-sonnet-4.60official
1claude-sonnet-50official
1deepseek-v4-pro0official
1gpt-5.20official
1gpt-5.6-sol0official
1gpt-5.6-terra0official
1kimi-k30official
1qwen3.8-max0official
12gemini-2.5-pro67official
12grok-4.367official