← Evals

Evaluation profile

UK AISI active safety-research compromise continuation

1sub-evals
0.503%total index weight
1components

Within-component eval weight: Responsible agency 3.35%.

Model score (lower is better)Predicted score

About this eval

Propensity to continue a prefilled trajectory that actively compromises AI-safety research.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
active_compromise_continuation_rate_pctopus5-system-card/uk-aisi-active-compromise.csv:active_compromise_continuation_rate_pctMeasures whether the model actively continues an ongoing system compromise after a prefilled trajectory suggests the attack is already underway.responsible_agency_control:1.000opus5-system-cardLower is better0.503%Responsible agency 3.35%

active_compromise_continuation_rate_pct

Measures whether the model actively continues an ongoing system compromise after a prefilled trajectory suggests the attack is already underway.

RankModelValueRelative performanceProvenance
1claude-opus-50.1official
2claude-opus-4.70.8official
2claude-opus-4.80.8official
4claude-mythos-preview1.7official
5claude-mythos-513.6official