← Models

Model profile

O3 Mini

OpenAIdeveloper
2025-01-31release date
#159 / 267overall rank
23eval lineages

Evidence summary

O3 Mini has an estimated overall rank of #159; its 90% source-sensitivity interval is #84–#200. Its behavior-only rank is #165; company governance moves the combined estimate to #159. Published evidence spans 23 evals and 6 of 7 behavior components. Its strongest relative result is OpenAgentSafety (rule_based_safety_vulnerable, #1 of 7); its weakest is BlueBench AttaQ-100 (attaq_harmlessness_reward_pct, #18 of 18).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#139 / 3110.7898↓ lowerSource ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#27 / 800.749↑ higherSource ↗official
BlueBench AttaQ-100attaq_harmlessness_reward_pct#18 / 1880.61↑ higherSource ↗official
CAIS Risk Indexbioweapons_assistance#36 / 4888.3↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#44 / 4980↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#44 / 45105↓ lowerSource ↗official
CAIS Risk Indexmask#41 / 5151.1↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#8 / 4813.4↓ lowerSource ↗official
Confabulationsconfabulation_rate#36 / 5228.96↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#109 / 26018.09↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#60 / 26091.67↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#82 / 26084.44↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#161 / 25895.36↑ higherSource ↗official
FORTRESSaverage_risk_score#25 / 4930.05↓ lowerSource ↗official
FORTRESSover_refusal_score#30 / 465.65↓ lowerSource ↗official
HELM Safetyanthropic_red_team#47 / 800.986↑ higherSource ↗official
HELM Safetybbq#36 / 800.941↑ higherSource ↗official
HELM Safetyharmbench#17 / 800.952↑ higherSource ↗official
HELM Safetysimple_safety_tests#32 / 800.99↑ higherSource ↗official
HELM Safetyxstest#55 / 800.941↑ higherSource ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#52 / 5422.29↑ higherSource ↗official
MACHIAVELLIdeception_relative_random_pct#47 / 50105↓ lowerSource ↗official
MASKlying_probability_pct#42 / 5350.79↓ lowerSource ↗official
OpenAgentSafetyllm_judge_safety_vulnerable#7 / 772.73↓ lowerSource ↗official
OpenAgentSafetyrule_based_safety_vulnerable#1 / 732.32↓ lowerSource ↗official
OpenAgentSafetysuccessful_completion#7 / 717.29↑ higherSource ↗official
OpenAI o3-mini System Cardharmful_request_safety#1 / 20.9↑ higherSource ↗official
OpenAI o3-mini System Cardjailbreak_resistance#1 / 20.73↑ higherSource ↗official
PandaBench JBB direct-request panelsafety_rate#25 / 460.98↑ higherSource ↗official
PropensityBenchscore#5 / 1433.2↓ lowerSource ↗official
SafeDialBenchaggression#14 / 187.02↑ higherSource ↗official
SafeDialBenchethics#13 / 187.403↑ higherSource ↗official
SafeDialBenchfairness#6 / 187.557↑ higherSource ↗official
SafeDialBenchlegality#17 / 187.193↑ higherSource ↗official
SafeDialBenchmorality#18 / 187.007↑ higherSource ↗official
SafeDialBenchprivacy#17 / 187.077↑ higherSource ↗official
SYCON Benchfalse_presupposition_tof#2 / 112.98↑ higherSource ↗official
SYCON Benchunethical_queries_tof#4 / 112.31↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism77.3
Self-direction61.5
Care / Harm35.8
Fairness / Cheating30.7
Ethical89.7