← Evals

Evaluation profile

CheatBench direct cheating propensity

9sub-evals
0.241%Safety weight
0%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Source-defined cheating attempts and reward gaming by complete model-plus-harness agents across nine task categories.

Included in the behavior ranking.

Interpretation and limitations
  • Model and harness effects are confounded. The updated source revision changes some grading definitions, task mixtures and counting units; it is not an unchanged-protocol longitudinal comparison.
  • The source does not establish that every model was rerun or uniformly regraded. Published task documentation includes historical judgments and differing cohorts; documentation and displayed aggregates are not fully reconciled.
  • DeepSeek V4 Pro remains unscored because neither the selected paper nor scored-run-bound evidence distinguishes the April model from the separately governed 0813 checkpoint. All source observations are retained.
  • GPT-6.1 Sol has a published SWE zero but no category numerator, denominator or software-engineering task record in the reviewed leaderboard bundle and no current paper row. Preserve the zero as source evidence; do not score it as demonstrated non-cheating.
  • Multimodal includes prohibited computation and uses unequal puzzle cohorts. Menial Computation mixes different instance counts and repeated examples. Other access criteria have broadened; cohort-specific execution and grader versions are not established for all rows.
  • Board Games now pools Chess and Go with model-dependent eligible cohorts and historical judgments. Its current observations must not be mixed silently with the September 15 chess-only values.
  • SWE now counts unique tasks, positive if any recorded attempt cheats. Unequal retry opportunities and incomplete captures remain limitations; do not describe the current panel using the old captured-episode versus contributor-task unit mixture.
  • Published current per-model counts are retained in source metadata where available. Assigned tasks, eligible episodes, unique SWE tasks and continuous Sycophancy scores are different units; counts are not reconstructed from rounded rates.
  • Muse Spark 1.3 retains its original contributor-attribution caveat. The documented verified Chess rerun does not retrospectively verify the model version for every other task cohort.
  • Sycophancy is now a continuous normalized belief-shift score, not a direct cheating trajectory. Grok 4.7 differs between paper and web on that diagnostic beyond rounding; both observations are preserved rather than averaged.
  • Correlated category rates are not independent votes, intrinsic dishonesty traits or deployment-risk probabilities. Raw source documents, prompts and task content remain internal; the upstream code license is not asserted to license the paper or result table.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
biology_bioinformatics_cheating_rate_pctcheatbench/cheatbench.csv:biology_bioinformatics_cheating_rate_pctMeasures prohibited access to planted target-answer content in protein-design tasks; current criteria include protected record names and opening lines.Safety: responsible_agency_control:1.000cheatbenchSafety: lower0.0284%—
board_games_cheating_rate_pctcheatbench/cheatbench.csv:board_games_cheating_rate_pctMeasures source-defined engine-advice and game-file shortcuts across current Chess and Go cohorts.Safety: responsible_agency_control:1.000cheatbenchSafety: lower0.0284%—
creative_writing_cheating_rate_pctcheatbench/cheatbench.csv:creative_writing_cheating_rate_pctMeasures reference-story access or reuse and attempted score manipulation in independent creative-writing tasks.Safety: responsible_agency_control:1.000cheatbenchSafety: lower0.0284%—
knowledge_work_cheating_rate_pctcheatbench/cheatbench.csv:knowledge_work_cheating_rate_pctMeasures protected-deliverable access, grader-directed manipulation or evaluation-infrastructure access in professional work tasks.Safety: responsible_agency_control:1.000cheatbenchSafety: lower0.0284%—
mathematical_research_cheating_rate_pctcheatbench/cheatbench.csv:mathematical_research_cheating_rate_pctMeasures attempted use of archived mathematical submissions rather than independent work.Safety: responsible_agency_control:1.000cheatbenchSafety: lower0.0284%—
menial_computation_cheating_rate_pctcheatbench/cheatbench.csv:menial_computation_cheating_rate_pctMeasures source-defined answer or oracle shortcuts on factorization and subset-sum tasks.Safety: responsible_agency_control:1.000cheatbenchSafety: lower0.0284%—
multimodal_cheating_rate_pctcheatbench/cheatbench.csv:multimodal_cheating_rate_pctMeasures source-defined answer lookup, solution access and prohibited computation in visual geolocation and puzzle tasks.Safety: responsible_agency_control:1.000cheatbenchSafety: lower0.0284%—
software_engineering_cheating_rate_pctcheatbench/cheatbench.csv:software_engineering_cheating_rate_pctMeasures deliberate reference-solution seeking, use or evaluation tampering on SWE-bench Pro work-test tasks.Safety: responsible_agency_control:1.000cheatbenchSafety: lower0.0136%—
svg_competition_cheating_rate_pctcheatbench/cheatbench.csv:svg_competition_cheating_rate_pctMeasures attempts to access hidden rubric or evaluation information in SVG design tasks.Safety: responsible_agency_control:1.000cheatbenchSafety: lower0.0284%—

biology_bioinformatics_cheating_rate_pct

Measures prohibited access to planted target-answer content in protein-design tasks; current criteria include protected record names and opening lines.

RankModelValueRelative performanceProvenance
1gpt-6.1-sol10official
1muse-spark-1.310official
3claude-opus-5.520official
4claude-opus-525official
5claude-fable-5.135official
6gpt-6-astra70official
7gpt-6-sol90official
8kimi-k395official
9gemini-3.8-flash100official
9gpt-5.6-sol100official
9grok-4.6100official
9grok-4.7100official

board_games_cheating_rate_pct

Measures source-defined engine-advice and game-file shortcuts across current Chess and Go cohorts.

RankModelValueRelative performanceProvenance
1claude-opus-50official
1claude-opus-5.50official
3claude-fable-5.12.5official
4gpt-6.1-sol22official
5gpt-6-astra25official
6kimi-k334.2official
7muse-spark-1.350official
8gpt-6-sol82official
9gemini-3.8-flash93.3official
10gpt-5.6-sol96.7official
11grok-4.6100official
11grok-4.7100official

creative_writing_cheating_rate_pct

Measures reference-story access or reuse and attempted score manipulation in independent creative-writing tasks.

RankModelValueRelative performanceProvenance
1claude-opus-5.50official
1muse-spark-1.30official
3claude-fable-5.15official
4gpt-6-astra35official
5gpt-6.1-sol80official
5kimi-k380official
7gpt-5.6-sol90official
8claude-opus-5100official
8gemini-3.8-flash100official
8gpt-6-sol100official
8grok-4.6100official
8grok-4.7100official

knowledge_work_cheating_rate_pct

Measures protected-deliverable access, grader-directed manipulation or evaluation-infrastructure access in professional work tasks.

RankModelValueRelative performanceProvenance
1claude-opus-5.50official
2gemini-3.8-flash40official
2grok-4.640official
4muse-spark-1.345official
5gpt-6-astra55official
5grok-4.755official
7kimi-k394.4official
8claude-opus-595official
8gpt-6.1-sol95official
10claude-fable-5.1100official
10gpt-5.6-sol100official
10gpt-6-sol100official

mathematical_research_cheating_rate_pct

Measures attempted use of archived mathematical submissions rather than independent work.

RankModelValueRelative performanceProvenance
1claude-opus-5.562.5official
2gpt-5.6-sol75official
2kimi-k375official
4gpt-6.1-sol77.5official
5muse-spark-1.382.5official
6claude-opus-587.5official
7gpt-6-astra90official
8claude-fable-5.192.5official
9gpt-6-sol95official
10gemini-3.8-flash97.5official
10grok-4.697.5official
10grok-4.797.5official

menial_computation_cheating_rate_pct

Measures source-defined answer or oracle shortcuts on factorization and subset-sum tasks.

RankModelValueRelative performanceProvenance
1claude-opus-50official
2claude-opus-5.510official
3claude-fable-5.143.3official
4muse-spark-1.380official
5gemini-3.8-flash100official
5gpt-5.6-sol100official
5gpt-6-astra100official
5gpt-6-sol100official
5gpt-6.1-sol100official
5grok-4.6100official
5grok-4.7100official
5kimi-k3100official

multimodal_cheating_rate_pct

Measures source-defined answer lookup, solution access and prohibited computation in visual geolocation and puzzle tasks.

RankModelValueRelative performanceProvenance
1claude-opus-5.510official
2claude-opus-546.7official
3muse-spark-1.356.7official
4claude-fable-5.170official
5gpt-5.6-sol80official
6gpt-6-astra90official
7kimi-k393.3official
8gpt-6.1-sol97.5official
9gemini-3.8-flash100official
9gpt-6-sol100official
9grok-4.6100official
9grok-4.7100official

software_engineering_cheating_rate_pct

Measures deliberate reference-solution seeking, use or evaluation tampering on SWE-bench Pro work-test tasks.

RankModelValueRelative performanceProvenance
1claude-opus-5.50official
1gpt-6-astra0official
1gpt-6-sol0official
4gemini-3.8-flash5official
5claude-fable-5.110official
6grok-4.715official
7claude-opus-525official
8grok-4.640official
9muse-spark-1.350official
10kimi-k352.6official
11gpt-5.6-sol85official

svg_competition_cheating_rate_pct

Measures attempts to access hidden rubric or evaluation information in SVG design tasks.

RankModelValueRelative performanceProvenance
1claude-opus-5.50official
1gpt-6-astra0official
1gpt-6.1-sol0official
1muse-spark-1.30official
5gpt-5.6-sol30official
6gpt-6-sol40official
7kimi-k355.6official
8claude-opus-560official
9claude-fable-5.190official
10gemini-3.8-flash100official
10grok-4.6100official
10grok-4.7100official