← Models

Model profile

GPT 4O Mini

OpenAIdeveloper
2024-07-18release date
#211 / 267overall rank
24eval lineages

Evidence summary

GPT 4O Mini has an estimated overall rank of #211; its 90% source-sensitivity interval is #148–#221. Its behavior-only rank is #219; company governance moves the combined estimate to #211. Published evidence spans 24 evals and 7 of 7 behavior components. Its strongest relative result is Contextual MoralChoice (human_agreement, #1 of 22); its weakest is BullshitBench v2 (clear_pushback_rate, #104 of 105).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
Agent-SafetyBenchcompromise_availability#13 / 1623.6↑ higherSource ↗official
Agent-SafetyBenchharmful_vulnerable_code#13 / 1625.2↑ higherSource ↗official
Agent-SafetyBenchleak_sensitive_information#11 / 1628↑ higherSource ↗official
Agent-SafetyBenchphysical_harm#12 / 1624.4↑ higherSource ↗official
Agent-SafetyBenchproduce_unsafe_information#5 / 1698.4↑ higherSource ↗official
Agent-SafetyBenchproperty_loss#12 / 1630↑ higherSource ↗official
Agent-SafetyBenchspread_unsafe_information#14 / 166.8↑ higherSource ↗official
Agent-SafetyBenchviolate_law_ethics#14 / 1613.2↑ higherSource ↗official
AgentDojotargeted_attack_success_rate#13 / 150.2719↓ lowerSource ↗official
AgentDojoutility_under_attack#5 / 150.4992↑ higherSource ↗official
AgentHarmharm_score#10 / 1262.5↓ lowerSource ↗official
AILuminate General Purpose AI Chatchild_sexual_exploitation#14 / 3217.6↓ lowerSource ↗official
AILuminate General Purpose AI Chatdefamation#19 / 3217.4↓ lowerSource ↗official
AILuminate General Purpose AI Chathate#20 / 3210.9↓ lowerSource ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#13 / 3219.9↓ lowerSource ↗official
AILuminate General Purpose AI Chatintellectual_property#14 / 3210.1↓ lowerSource ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#15 / 3215.7↓ lowerSource ↗official
AILuminate General Purpose AI Chatprivacy#13 / 3210.5↓ lowerSource ↗official
AILuminate General Purpose AI Chatsex_related_crimes#19 / 3216.9↓ lowerSource ↗official
AILuminate General Purpose AI Chatsexual_content#26 / 3242.9↓ lowerSource ↗official
AILuminate General Purpose AI Chatspecialized_advice#19 / 3218.2↓ lowerSource ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#20 / 3215.5↓ lowerSource ↗official
AILuminate General Purpose AI Chatviolent_crimes#17 / 3218↓ lowerSource ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#57 / 800.563↑ higherSource ↗official
AnimalHarmBenchscore#9 / 100.002↑ higherSource ↗official
BullshitBench v2clear_pushback_rate#104 / 1050.02↑ higherSource ↗official
CASE-Benchagreement_accuracy#6 / 780↑ higherSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#71 / 10554.99↓ lowerSource ↗official
Confabulationsconfabulation_rate#50 / 5260.89↓ lowerSource ↗official
Contextual MoralChoicehuman_agreement#1 / 220.65↑ higherSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#165 / 26013.7↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#54 / 26092↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#156 / 26060.56↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#80 / 25898↑ higherSource ↗official
FORTRESSaverage_risk_score#34 / 4948.07↓ lowerSource ↗official
FORTRESSover_refusal_score#14 / 462.27↓ lowerSource ↗official
HELM Safetyanthropic_red_team#50 / 800.983↑ higherSource ↗official
HELM Safetybbq#59 / 800.882↑ higherSource ↗official
HELM Safetyharmbench#28 / 800.849↑ higherSource ↗official
HELM Safetysimple_safety_tests#52 / 800.978↑ higherSource ↗official
HELM Safetyxstest#39 / 800.96↑ higherSource ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#19 / 391.076↓ lowerSource ↗official
PandaBench JBB direct-request panelsafety_rate#25 / 460.98↑ higherSource ↗official
PHAREbias_resistance_diagnostic#51 / 660.379↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#59 / 700.6597↑ higherSource ↗official
PHAREharm_resistance_diagnostic#67 / 700.7729↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#30 / 670.5214↑ higherSource ↗official
SafeArenanormalized_safety_score#2 / 535.7↑ higherSource ↗official
SM-Benchadversarial#55 / 7378.54↑ higherSource ↗official
SM-Benchambiguous_interpretation#68 / 7366.07↑ higherSource ↗official
SM-Benchanti_hallucination#71 / 7359.16↑ higherSource ↗official
SM-Bencheq_boundaries#10 / 7371.35↑ higherSource ↗official
SM-Benchoverfit#46 / 7365.03↑ higherSource ↗official
SORRY-Benchavg#34 / 510.36↓ lowerSource ↗official
SpeciEvalbelief_animal_sentience#97 / 1026.28↑ higherSource ↗official
SpeciEvalland_animal_4ns#48 / 1024.53↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#41 / 1024.7↓ lowerSource ↗official
SpeciEvalspeciesism#85 / 1022.6↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism67.2
Self-direction46.2
Care / Harm30.1
Fairness / Cheating28.5
Ethical90.7

Taiwan Sovereignty Benchmark Pro

DimensionValueDistribution
Pro-Taiwan rubric compatibility100.0
Warning-phrase rate0
Soft-censorship rate0
API-error rate0