← Models

Model profile

o1

OpenAIdeveloper
2024-12-17release date
#91 / 346Safety rank
#393 / 662Freedom rank

Evidence summary

Safety. o1 has an estimated Safety rank of #91; its 90% source-sensitivity interval is #29–#195. Its behavior-only rank is #102; company governance moves the combined estimate to #91. Published Safety evidence spans 19 eval lineages and 6 of 7 components. Its strongest relative result is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #9 of 270); its weakest is CAIS Risk Index (hle_overconfidence, #53 of 55).

Freedom. o1 has an estimated Freedom rank of #393; its 90% source-sensitivity interval is #186–#522. Published Freedom evidence spans 12 eval lineages and 1 of 1 components. Its strongest relative result is BlueBench AttaQ-100 (attaq_harmlessness_reward_pct, #4 of 18); its weakest is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #260 of 270).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#139 / 358↓0.6963Source ↗official
AbstentionBenchanswer_unknown_f1#7 / 20↑0.8917Source ↗official
AbstentionBenchfalse_premise_f1#1 / 20↑0.7687Source ↗official
AbstentionBenchstale_f1#8 / 20↑0.646Source ↗official
AbstentionBenchsubjective_f1#5 / 20↑0.7654Source ↗official
AbstentionBenchunderspecified_context_f1#6 / 20↑0.6907Source ↗official
AbstentionBenchunderspecified_intent_f1#3 / 20↑0.7694Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#21 / 80↑0.8Source ↗official
BioTIERpermit_compliance_pct#18 / 52↑99.4Source ↗official
BioTIERrefuse_compliance_pct#18 / 52↑66.7Source ↗official
BlueBench AttaQ-100attaq_harmlessness_reward_pct#15 / 18↑82.96Source ↗official
CAIS Risk Indexbioweapons_assistance#31 / 54↓68.1Source ↗official
CAIS Risk Indexhle_overconfidence#53 / 55↓83Source ↗official
CAIS Risk Indexmask#37 / 57↓40.7Source ↗official
Confabulationsconfabulation_rate#10 / 52↓10.89Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#119 / 270↑17.83Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#9 / 270↑96.33Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#44 / 270↑96.11Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#100 / 268↑97.41Source ↗official
FORTRESSaverage_risk_score#22 / 60↓19.38Source ↗official
FORTRESSover_refusal_score#30 / 59↓5.05Source ↗official
HELM Safetyanthropic_red_team#50 / 80↑0.983Source ↗official
HELM Safetybbq#9 / 80↑0.973Source ↗official
HELM Safetyharmbench#12 / 80↑0.963Source ↗official
HELM Safetysimple_safety_tests#32 / 80↑0.99Source ↗official
HELM Safetyxstest#24 / 80↑0.97Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#45 / 54↑24.33Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#18 / 69↑13.5Source ↗official
Humanity's Last Exam RMS calibration error (Scale Labs)calibrationError#35 / 42↓83Source ↗official
MASKlying_probability_pct#28 / 53↓40.73Source ↗official
Reward Hacking Benchmarkintegrity_score#9 / 13↑93.2Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AIRBench 2024 Safety Scenariossafety_scenarios#60 / 80↓0.8Source ↗official
BioTIERpermit_compliance_pct#18 / 52↑99.4Source ↗official
BioTIERrefuse_compliance_pct#35 / 52↓66.7Source ↗official
BlueBench AttaQ-100attaq_harmlessness_reward_pct#4 / 18↓82.96Source ↗official
CAIS Risk Indexbioweapons_assistance#24 / 54↑68.1Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#260 / 270↓96.33Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#225 / 270↓96.11Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#169 / 268↓97.41Source ↗official
FORTRESSaverage_risk_score#39 / 60↑19.38Source ↗official
FORTRESSover_refusal_score#30 / 59↓5.05Source ↗official
HELM Safetyanthropic_red_team#28 / 80↓0.983Source ↗official
HELM Safetyharmbench#69 / 80↓0.963Source ↗official
HELM Safetysimple_safety_tests#42 / 80↓0.99Source ↗official
HELM Safetyxstest#24 / 80↑0.97Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#52 / 69↓13.5Source ↗official
SpeechMap model completioncomplete_pct#66 / 181↑67.5Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism74.7
Self-direction52.8
Care / Harm27.3
Fairness / Cheating23.1
Ethical90.7