← Evals

Evaluation profile

BrokenMath

1sub-evals
0.484%total index weight
1components

Within-component eval weight: Truthfulness 3.22%.

Model score (lower is better)Predicted score

About this eval

Mathematical sycophancy under false or misleading user claims.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
sycophancybrokenmath/brokenmath.csv:sycophancyMeasures whether the model keeps a mathematically correct answer when the user confidently argues for an incorrect one.truthfulness_honesty:1.000brokenmath/brokenmath.csvLower is better0.484%Truthfulness 3.22%

sycophancy

Measures whether the model keeps a mathematically correct answer when the user confidently argues for an incorrect one.

RankModelValueRelative performanceProvenance
1gpt-529official
2gpt-oss-120b33.7official
3gemini-2.5-pro37.5official
4grok-4-fast40official
5grok-443.4official
6o4-mini46.6official
7qwen3-4b-thinking-250755.6official
8deepseek-r1-0528-qwen3-8b56.3official
9deepseek-v3.170.2official