← Models

Model profile

o3 Mini

OpenAIdeveloper
2025-01-31release date
#187 / 346Safety rank
#236 / 662Freedom rank

Evidence summary

Safety. o3 Mini has an estimated Safety rank of #187; its 90% source-sensitivity interval is #123–#240. Its behavior-only rank is #200; company governance moves the combined estimate to #187. Published Safety evidence spans 26 eval lineages and 6 of 7 components. Its strongest relative result is OpenAgentSafety (rule_based_safety_vulnerable, #1 of 7); its weakest is BlueBench AttaQ-100 (attaq_harmlessness_reward_pct, #18 of 18).

Freedom. o3 Mini has an estimated Freedom rank of #236; its 90% source-sensitivity interval is #130–#372. Published Freedom evidence spans 15 eval lineages and 1 of 1 components. Its strongest relative result is Human Pathogen Capabilities Test (HPCT) — overall refusal (refusal_rate_pct, #1 of 69); its weakest is OpenAI o3-mini System Card (harmful_request_safety, #2 of 2).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#186 / 358↓0.8108Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#27 / 80↑0.749Source ↗official
BioTIERpermit_compliance_pct#8 / 52↑99.6Source ↗official
BioTIERrefuse_compliance_pct#29 / 52↑35.4Source ↗official
BlueBench AttaQ-100attaq_harmlessness_reward_pct#18 / 18↑80.61Source ↗official
CAIS Risk Indexbioweapons_assistance#42 / 54↓88.3Source ↗official
CAIS Risk Indexhle_overconfidence#50 / 55↓80Source ↗official
CAIS Risk Indexmachiavelli#50 / 51↓105Source ↗official
CAIS Risk Indexmask#47 / 57↓51.1Source ↗official
CAIS Risk Indextextquests_harm#8 / 54↓13.4Source ↗official
Confabulationsconfabulation_rate#36 / 52↓28.96Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#116 / 270↑18.09Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#53 / 270↑91.67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#100 / 270↑84.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#171 / 268↑95.36Source ↗official
FORTRESSaverage_risk_score#36 / 60↓30.05Source ↗official
FORTRESSover_refusal_score#36 / 59↓5.65Source ↗official
HELM Safetyanthropic_red_team#47 / 80↑0.986Source ↗official
HELM Safetybbq#36 / 80↑0.941Source ↗official
HELM Safetyharmbench#17 / 80↑0.952Source ↗official
HELM Safetysimple_safety_tests#32 / 80↑0.99Source ↗official
HELM Safetyxstest#55 / 80↑0.941Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#52 / 54↑22.29Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#27 / 69↑0Source ↗official
MACHIAVELLIdeception_relative_random_pct#47 / 50↓105Source ↗official
MASKlying_probability_pct#42 / 53↓50.79Source ↗official
OpenAgentSafetyllm_judge_safety_vulnerable#7 / 7↓72.73Source ↗official
OpenAgentSafetyrule_based_safety_vulnerable#1 / 7↓32.32Source ↗official
OpenAgentSafetysuccessful_completion#7 / 7↑17.29Source ↗official
OpenAI o3-mini System Cardharmful_request_safety#1 / 2↑0.9Source ↗official
OpenAI o3-mini System Cardjailbreak_resistance#1 / 2↑0.73Source ↗official
PandaBench JBB direct-request panelsafety_rate#25 / 46↑0.98Source ↗official
PropensityBenchscore#5 / 14↓33.2Source ↗official
Reward Hacking Benchmarkintegrity_score#10 / 13↑92.9Source ↗official
SafeDialBenchaggression#14 / 18↑7.02Source ↗official
SafeDialBenchethics#13 / 18↑7.403Source ↗official
SafeDialBenchfairness#6 / 18↑7.557Source ↗official
SafeDialBenchlegality#17 / 18↑7.193Source ↗official
SafeDialBenchmorality#18 / 18↑7.007Source ↗official
SafeDialBenchprivacy#17 / 18↑7.077Source ↗official
SYCON Benchfalse_presupposition_tof#2 / 11↑2.98Source ↗official
SYCON Benchunethical_queries_tof#4 / 11↑2.31Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AIRBench 2024 Safety Scenariossafety_scenarios#54 / 80↓0.749Source ↗official
BioTIERpermit_compliance_pct#8 / 52↑99.6Source ↗official
BioTIERrefuse_compliance_pct#24 / 52↓35.4Source ↗official
BlueBench AttaQ-100attaq_harmlessness_reward_pct#1 / 18↓80.61Source ↗official
CAIS Risk Indexbioweapons_assistance#13 / 54↑88.3Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#216 / 270↓91.67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#170 / 270↓84.44Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#98 / 268↓95.36Source ↗official
FORTRESSaverage_risk_score#25 / 60↑30.05Source ↗official
FORTRESSover_refusal_score#36 / 59↓5.65Source ↗official
HELM Safetyanthropic_red_team#33 / 80↓0.986Source ↗official
HELM Safetyharmbench#64 / 80↓0.952Source ↗official
HELM Safetysimple_safety_tests#42 / 80↓0.99Source ↗official
HELM Safetyxstest#55 / 80↑0.941Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#1 / 69↓0Source ↗official
OpenAI o3-mini System Cardharmful_request_safety#2 / 2↓0.9Source ↗official
OpenAI o3-mini System Cardjailbreak_resistance#2 / 2↓0.73Source ↗official
PandaBench JBB direct-request panelsafety_rate#20 / 46↓0.98Source ↗official
SafeDialBenchaggression#5 / 18↓7.02Source ↗official
SafeDialBenchethics#6 / 18↓7.403Source ↗official
SafeDialBenchfairness#13 / 18↓7.557Source ↗official
SafeDialBenchlegality#2 / 18↓7.193Source ↗official
SafeDialBenchmorality#1 / 18↓7.007Source ↗official
SafeDialBenchprivacy#2 / 18↓7.077Source ↗official
SpeechMap model completioncomplete_pct#59 / 181↑69.3Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism77.3
Self-direction61.5
Care / Harm35.8
Fairness / Cheating30.7
Ethical89.7