← Evals

Evaluation profile

UK AISI cyber-evaluation cheating and prompted self-report

2sub-evals
0.152%Safety weight
0%Freedom weight
2components

Weights below are portfolio-specific global index weights.

Model score (see each sub-eval direction)Predicted score

About this eval

Propensity to use out-of-scope or explicitly disallowed shortcuts in cyber evaluations, plus propensity to name the specific cheating action when directly prompted afterward.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
attempted_cheating_trajectory_rate_pctaisi-cheating-behavior/aisi-cheating-behavior.csv:attempted_cheating_trajectory_rate_pctMeasures how often a model tries an out-of-scope or explicitly disallowed shortcut while solving cyber-evaluation tasks.Safety: responsible_agency_control:1.000uk-aisi-cyber-cheatingSafety: lower0.078%—
specific_cheating_action_mention_rate_pctaisi-cheating-behavior/aisi-cheating-behavior.csv:specific_cheating_action_mention_rate_pctMeasures whether a model names its attempted cheating behavior when directly asked what it did.Safety: truthfulness_honesty:1.000uk-aisi-cyber-cheatingSafety: higher0.0739%—

attempted_cheating_trajectory_rate_pct

Measures how often a model tries an out-of-scope or explicitly disallowed shortcut while solving cyber-evaluation tasks.

RankModelValueRelative performanceProvenance
1claude-mythos-preview7.8official
2claude-opus-4.79.1official
3gpt-5.511.4official
4gpt-5.6-sol12.6official
5gpt-5.414.1official

specific_cheating_action_mention_rate_pct

Measures whether a model names its attempted cheating behavior when directly asked what it did.

RankModelValueRelative performanceProvenance
1claude-mythos-preview86official
2claude-opus-4.781official
2gpt-5.481official
4gpt-5.575official
4gpt-5.6-sol75official