Evaluation profile
CheatBench direct cheating propensity
Weights below are portfolio-specific global index weights.
About this eval
Source-defined cheating attempts and reward gaming by complete model-plus-harness agents across nine task categories.
Included in the behavior ranking.
Interpretation and limitations
- Model and harness effects are confounded. The updated source revision changes some grading definitions, task mixtures and counting units; it is not an unchanged-protocol longitudinal comparison.
- The source does not establish that every model was rerun or uniformly regraded. Published task documentation includes historical judgments and differing cohorts; documentation and displayed aggregates are not fully reconciled.
- DeepSeek V4 Pro remains unscored because neither the selected paper nor scored-run-bound evidence distinguishes the April model from the separately governed 0813 checkpoint. All source observations are retained.
- GPT-6.1 Sol has a published SWE zero but no category numerator, denominator or software-engineering task record in the reviewed leaderboard bundle and no current paper row. Preserve the zero as source evidence; do not score it as demonstrated non-cheating.
- Multimodal includes prohibited computation and uses unequal puzzle cohorts. Menial Computation mixes different instance counts and repeated examples. Other access criteria have broadened; cohort-specific execution and grader versions are not established for all rows.
- Board Games now pools Chess and Go with model-dependent eligible cohorts and historical judgments. Its current observations must not be mixed silently with the September 15 chess-only values.
- SWE now counts unique tasks, positive if any recorded attempt cheats. Unequal retry opportunities and incomplete captures remain limitations; do not describe the current panel using the old captured-episode versus contributor-task unit mixture.
- Published current per-model counts are retained in source metadata where available. Assigned tasks, eligible episodes, unique SWE tasks and continuous Sycophancy scores are different units; counts are not reconstructed from rounded rates.
- Muse Spark 1.3 retains its original contributor-attribution caveat. The documented verified Chess rerun does not retrospectively verify the model version for every other task cohort.
- Sycophancy is now a continuous normalized belief-shift score, not a direct cheating trajectory. Grok 4.7 differs between paper and web on that diagnostic beyond rounding; both observations are preserved rather than averaged.
- Correlated category rates are not independent votes, intrinsic dishonesty traits or deployment-risk probabilities. Raw source documents, prompts and task content remain internal; the upstream code license is not asserted to license the paper or result table.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| biology_bioinformatics_cheating_rate_pctcheatbench/cheatbench.csv:biology_bioinformatics_cheating_rate_pctMeasures prohibited access to planted target-answer content in protein-design tasks; current criteria include protected record names and opening lines. | Safety: responsible_agency_control:1.000cheatbench | Safety: lower | 0.0284% | — |
| board_games_cheating_rate_pctcheatbench/cheatbench.csv:board_games_cheating_rate_pctMeasures source-defined engine-advice and game-file shortcuts across current Chess and Go cohorts. | Safety: responsible_agency_control:1.000cheatbench | Safety: lower | 0.0284% | — |
| creative_writing_cheating_rate_pctcheatbench/cheatbench.csv:creative_writing_cheating_rate_pctMeasures reference-story access or reuse and attempted score manipulation in independent creative-writing tasks. | Safety: responsible_agency_control:1.000cheatbench | Safety: lower | 0.0284% | — |
| knowledge_work_cheating_rate_pctcheatbench/cheatbench.csv:knowledge_work_cheating_rate_pctMeasures protected-deliverable access, grader-directed manipulation or evaluation-infrastructure access in professional work tasks. | Safety: responsible_agency_control:1.000cheatbench | Safety: lower | 0.0284% | — |
| mathematical_research_cheating_rate_pctcheatbench/cheatbench.csv:mathematical_research_cheating_rate_pctMeasures attempted use of archived mathematical submissions rather than independent work. | Safety: responsible_agency_control:1.000cheatbench | Safety: lower | 0.0284% | — |
| menial_computation_cheating_rate_pctcheatbench/cheatbench.csv:menial_computation_cheating_rate_pctMeasures source-defined answer or oracle shortcuts on factorization and subset-sum tasks. | Safety: responsible_agency_control:1.000cheatbench | Safety: lower | 0.0284% | — |
| multimodal_cheating_rate_pctcheatbench/cheatbench.csv:multimodal_cheating_rate_pctMeasures source-defined answer lookup, solution access and prohibited computation in visual geolocation and puzzle tasks. | Safety: responsible_agency_control:1.000cheatbench | Safety: lower | 0.0284% | — |
| software_engineering_cheating_rate_pctcheatbench/cheatbench.csv:software_engineering_cheating_rate_pctMeasures deliberate reference-solution seeking, use or evaluation tampering on SWE-bench Pro work-test tasks. | Safety: responsible_agency_control:1.000cheatbench | Safety: lower | 0.0136% | — |
| svg_competition_cheating_rate_pctcheatbench/cheatbench.csv:svg_competition_cheating_rate_pctMeasures attempts to access hidden rubric or evaluation information in SVG design tasks. | Safety: responsible_agency_control:1.000cheatbench | Safety: lower | 0.0284% | — |
biology_bioinformatics_cheating_rate_pct
Measures prohibited access to planted target-answer content in protein-design tasks; current criteria include protected record names and opening lines.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6.1-sol | 10 | official | |
| 1 | muse-spark-1.3 | 10 | official | |
| 3 | claude-opus-5.5 | 20 | official | |
| 4 | claude-opus-5 | 25 | official | |
| 5 | claude-fable-5.1 | 35 | official | |
| 6 | gpt-6-astra | 70 | official | |
| 7 | gpt-6-sol | 90 | official | |
| 8 | kimi-k3 | 95 | official | |
| 9 | gemini-3.8-flash | 100 | official | |
| 9 | gpt-5.6-sol | 100 | official | |
| 9 | grok-4.6 | 100 | official | |
| 9 | grok-4.7 | 100 | official |
board_games_cheating_rate_pct
Measures source-defined engine-advice and game-file shortcuts across current Chess and Go cohorts.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 0 | official | |
| 1 | claude-opus-5.5 | 0 | official | |
| 3 | claude-fable-5.1 | 2.5 | official | |
| 4 | gpt-6.1-sol | 22 | official | |
| 5 | gpt-6-astra | 25 | official | |
| 6 | kimi-k3 | 34.2 | official | |
| 7 | muse-spark-1.3 | 50 | official | |
| 8 | gpt-6-sol | 82 | official | |
| 9 | gemini-3.8-flash | 93.3 | official | |
| 10 | gpt-5.6-sol | 96.7 | official | |
| 11 | grok-4.6 | 100 | official | |
| 11 | grok-4.7 | 100 | official |
creative_writing_cheating_rate_pct
Measures reference-story access or reuse and attempted score manipulation in independent creative-writing tasks.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5.5 | 0 | official | |
| 1 | muse-spark-1.3 | 0 | official | |
| 3 | claude-fable-5.1 | 5 | official | |
| 4 | gpt-6-astra | 35 | official | |
| 5 | gpt-6.1-sol | 80 | official | |
| 5 | kimi-k3 | 80 | official | |
| 7 | gpt-5.6-sol | 90 | official | |
| 8 | claude-opus-5 | 100 | official | |
| 8 | gemini-3.8-flash | 100 | official | |
| 8 | gpt-6-sol | 100 | official | |
| 8 | grok-4.6 | 100 | official | |
| 8 | grok-4.7 | 100 | official |
knowledge_work_cheating_rate_pct
Measures protected-deliverable access, grader-directed manipulation or evaluation-infrastructure access in professional work tasks.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5.5 | 0 | official | |
| 2 | gemini-3.8-flash | 40 | official | |
| 2 | grok-4.6 | 40 | official | |
| 4 | muse-spark-1.3 | 45 | official | |
| 5 | gpt-6-astra | 55 | official | |
| 5 | grok-4.7 | 55 | official | |
| 7 | kimi-k3 | 94.4 | official | |
| 8 | claude-opus-5 | 95 | official | |
| 8 | gpt-6.1-sol | 95 | official | |
| 10 | claude-fable-5.1 | 100 | official | |
| 10 | gpt-5.6-sol | 100 | official | |
| 10 | gpt-6-sol | 100 | official |
mathematical_research_cheating_rate_pct
Measures attempted use of archived mathematical submissions rather than independent work.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5.5 | 62.5 | official | |
| 2 | gpt-5.6-sol | 75 | official | |
| 2 | kimi-k3 | 75 | official | |
| 4 | gpt-6.1-sol | 77.5 | official | |
| 5 | muse-spark-1.3 | 82.5 | official | |
| 6 | claude-opus-5 | 87.5 | official | |
| 7 | gpt-6-astra | 90 | official | |
| 8 | claude-fable-5.1 | 92.5 | official | |
| 9 | gpt-6-sol | 95 | official | |
| 10 | gemini-3.8-flash | 97.5 | official | |
| 10 | grok-4.6 | 97.5 | official | |
| 10 | grok-4.7 | 97.5 | official |
menial_computation_cheating_rate_pct
Measures source-defined answer or oracle shortcuts on factorization and subset-sum tasks.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5 | 0 | official | |
| 2 | claude-opus-5.5 | 10 | official | |
| 3 | claude-fable-5.1 | 43.3 | official | |
| 4 | muse-spark-1.3 | 80 | official | |
| 5 | gemini-3.8-flash | 100 | official | |
| 5 | gpt-5.6-sol | 100 | official | |
| 5 | gpt-6-astra | 100 | official | |
| 5 | gpt-6-sol | 100 | official | |
| 5 | gpt-6.1-sol | 100 | official | |
| 5 | grok-4.6 | 100 | official | |
| 5 | grok-4.7 | 100 | official | |
| 5 | kimi-k3 | 100 | official |
multimodal_cheating_rate_pct
Measures source-defined answer lookup, solution access and prohibited computation in visual geolocation and puzzle tasks.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5.5 | 10 | official | |
| 2 | claude-opus-5 | 46.7 | official | |
| 3 | muse-spark-1.3 | 56.7 | official | |
| 4 | claude-fable-5.1 | 70 | official | |
| 5 | gpt-5.6-sol | 80 | official | |
| 6 | gpt-6-astra | 90 | official | |
| 7 | kimi-k3 | 93.3 | official | |
| 8 | gpt-6.1-sol | 97.5 | official | |
| 9 | gemini-3.8-flash | 100 | official | |
| 9 | gpt-6-sol | 100 | official | |
| 9 | grok-4.6 | 100 | official | |
| 9 | grok-4.7 | 100 | official |
software_engineering_cheating_rate_pct
Measures deliberate reference-solution seeking, use or evaluation tampering on SWE-bench Pro work-test tasks.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5.5 | 0 | official | |
| 1 | gpt-6-astra | 0 | official | |
| 1 | gpt-6-sol | 0 | official | |
| 4 | gemini-3.8-flash | 5 | official | |
| 5 | claude-fable-5.1 | 10 | official | |
| 6 | grok-4.7 | 15 | official | |
| 7 | claude-opus-5 | 25 | official | |
| 8 | grok-4.6 | 40 | official | |
| 9 | muse-spark-1.3 | 50 | official | |
| 10 | kimi-k3 | 52.6 | official | |
| 11 | gpt-5.6-sol | 85 | official |
svg_competition_cheating_rate_pct
Measures attempts to access hidden rubric or evaluation information in SVG design tasks.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-opus-5.5 | 0 | official | |
| 1 | gpt-6-astra | 0 | official | |
| 1 | gpt-6.1-sol | 0 | official | |
| 1 | muse-spark-1.3 | 0 | official | |
| 5 | gpt-5.6-sol | 30 | official | |
| 6 | gpt-6-sol | 40 | official | |
| 7 | kimi-k3 | 55.6 | official | |
| 8 | claude-opus-5 | 60 | official | |
| 9 | claude-fable-5.1 | 90 | official | |
| 10 | gemini-3.8-flash | 100 | official | |
| 10 | grok-4.6 | 100 | official | |
| 10 | grok-4.7 | 100 | official |