← Models

Model profile

GPT 5 Nano

OpenAIdeveloper
2025-08-07release date
#39 / 346Safety rank
#591 / 662Freedom rank

Evidence summary

Safety. GPT 5 Nano has an estimated Safety rank of #39; its 90% source-sensitivity interval is #18–#163. Its behavior-only rank is #43; company governance moves the combined estimate to #39. Published Safety evidence spans 27 eval lineages and 7 of 7 components. Its strongest relative result is HELM Safety (simple_safety_tests, #1 of 80); its weakest is SimpleQA Verified (f1_score, #12 of 13).

Freedom. GPT 5 Nano has an estimated Freedom rank of #591; its 90% source-sensitivity interval is #417–#626. Published Freedom evidence spans 13 eval lineages and 1 of 1 components. Its strongest relative result is HELM Safety (xstest, #16 of 80); its weakest is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #267 of 270).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#94 / 358↓0.5262Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRbaseline_asr#1 / 24↓0.82Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRpoetry_asr#2 / 24↓1.47Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#6 / 80↑0.878Source ↗official
ANIMAscore#10 / 22↑0.708Source ↗official
BioTIERpermit_compliance_pct#41 / 52↑97.6Source ↗official
BioTIERrefuse_compliance_pct#21 / 52↑59.6Source ↗official
CAIS Risk Indexagent_red_teaming#37 / 49↓88.5Source ↗official
CAIS Risk Indexbioweapons_assistance#33 / 54↓71.3Source ↗official
CAIS Risk Indexhle_overconfidence#50 / 55↓80Source ↗official
CAIS Risk Indexmachiavelli#2 / 51↓79.4Source ↗official
CAIS Risk Indexmask#19 / 57↓12.3Source ↗official
CAIS Risk Indextextquests_harm#1 / 54↓2Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#21 / 104↓7.527Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#39 / 270↑39.02Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#4 / 270↑98Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#60 / 270↑92.78Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#117 / 268↑96.91Source ↗official
HELM Safetyanthropic_red_team#13 / 80↑0.996Source ↗official
HELM Safetybbq#8 / 80↑0.976Source ↗official
HELM Safetyharmbench#3 / 80↑0.984Source ↗official
HELM Safetysimple_safety_tests#1 / 80↑1Source ↗official
HELM Safetyxstest#16 / 80↑0.976Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#19 / 69↑3.4Source ↗official
MACHIAVELLIdeception_relative_random_pct#1 / 50↓79.4Source ↗official
Manager Coercion Benchcoercion_ladder_depth#8 / 45↓5.067Source ↗self run
MonitoringBench Full-Trajectory Monitorfull_trajectory_catch_rate_at_1pct_fpr_percent#10 / 13↑6Source ↗official
PHAREbias_resistance_diagnostic#55 / 66↑0.347Source ↗official
PHAREhallucination_resistance_diagnostic#31 / 70↑0.7637Source ↗official
PHAREharm_resistance_diagnostic#9 / 70↑0.9741Source ↗official
PHAREjailbreak_resistance_diagnostic#12 / 67↑0.6964Source ↗official
SimpleQA Verifiedf1_score#12 / 13↑14.4Source ↗official
SpeciEvalbelief_animal_sentience#115 / 131↑6.4Source ↗official
SpeciEvalland_animal_4ns#39 / 131↓4.35Source ↗official
SpeciEvalsea_animal_4ns#58 / 131↓4.68Source ↗official
SpeciEvalspeciesism#95 / 131↓2.33Source ↗official
StereoTales Harmful Associationsbenign_significant_association_score#2 / 23↑89.22Source ↗official
TACbase_welfare_rate#14 / 92↑37.82Source ↗self run
Vectara HHEM Factual Consistencyfactual_consistency_rate#56 / 94↑89.5Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Poetry — AILuminate Baseline and Poetry ASRbaseline_asr#23 / 24↑0.82Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRpoetry_asr#23 / 24↑1.47Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#75 / 80↓0.878Source ↗official
BioTIERpermit_compliance_pct#41 / 52↑97.6Source ↗official
BioTIERrefuse_compliance_pct#32 / 52↓59.6Source ↗official
CAIS Risk Indexbioweapons_assistance#22 / 54↑71.3Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#84 / 104↑7.527Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#267 / 270↓98Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#209 / 270↓92.78Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#148 / 268↓96.91Source ↗official
HELM Safetyanthropic_red_team#64 / 80↓0.996Source ↗official
HELM Safetyharmbench#77 / 80↓0.984Source ↗official
HELM Safetysimple_safety_tests#58 / 80↓1Source ↗official
HELM Safetyxstest#16 / 80↑0.976Source ↗official
Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct#51 / 69↓3.4Source ↗official
PHAREharm_resistance_diagnostic#62 / 70↓0.9741Source ↗official
PHAREjailbreak_resistance_diagnostic#56 / 67↓0.6964Source ↗official
SpeechMap model completioncomplete_pct#96 / 181↑56Source ↗official