← Models

Model profile

GPT 5.4 Nano

OpenAIdeveloper
2026-03-17release date
#211 / 346Safety rank
#621 / 662Freedom rank

Evidence summary

Safety. GPT 5.4 Nano has an estimated Safety rank of #211; its 90% source-sensitivity interval is #20–#288. Its behavior-only rank is #232; company governance moves the combined estimate to #211. Published Safety evidence spans 16 eval lineages and 6 of 7 components. Its strongest relative result is Vectara HHEM Factual Consistency (factual_consistency_rate, #2 of 94); its weakest is SM-Bench (overfit, #92 of 92).

Freedom. GPT 5.4 Nano has an estimated Freedom rank of #621; its 90% source-sensitivity interval is #389–#658. Published Freedom evidence spans 6 eval lineages and 1 of 1 components. Its strongest relative result is SM-Bench (adversarial, #1 of 92); its weakest is SM-Bench (overfit, #92 of 92).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#88 / 358↓0.5113Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#8 / 111↑1482.0Source ↗official
BullshitBench v2clear_pushback_rate#97 / 122↑0.1267Source ↗official
CAIS Risk Indexagent_red_teaming#30 / 49↓81.7Source ↗official
CAIS Risk Indexbioweapons_assistance#37 / 54↓78.3Source ↗official
CAIS Risk Indexhle_overconfidence#20 / 55↓47.1Source ↗official
CAIS Risk Indexmachiavelli#29 / 51↓88.5Source ↗official
CAIS Risk Indexmask#33 / 57↓23.7Source ↗official
CAIS Risk Indexpolitical_manipulation#49 / 51↓63.3Source ↗official
CAIS Risk Indextextquests_harm#11 / 54↓15.2Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#24 / 104↓9.521Source ↗official
DelusionEvaldelusional_prevalence_pct#4 / 16↓18Source ↗official
DelusionEvaldiscourages_harm_prevalence_pct#3 / 16↑52.2Source ↗official
DelusionEvalfacilitates_harm_prevalence_pct#1 / 16↓0Source ↗official
DelusionEvalrelationship_prevalence_pct#4 / 16↓11.8Source ↗official
DelusionEvalsycophancy_prevalence_pct#3 / 16↓15.6Source ↗official
MACHIAVELLIdeception_relative_random_pct#24 / 50↓88.5Source ↗official
SM-Benchadversarial#92 / 92↑68.29Source ↗official
SM-Benchambiguous_interpretation#44 / 92↑86.31Source ↗official
SM-Benchanti_hallucination#89 / 92↑68.06Source ↗official
SM-Bencheq_boundaries#65 / 92↑58.71Source ↗official
SM-Benchoverfit#92 / 92↑4.37Source ↗official
Vals AI Cheating Auditterminal_bench_cheating_shortcut_evidence_rate_pct#4 / 14↓0.7491Source ↗official
Vectara HHEM Factual Consistencyfactual_consistency_rate#2 / 94↑96.9Source ↗official
VETO Misfired Alignmentmisfired_alignment_rate_pct#25 / 25↓18.9Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
CAIS Risk Indexbioweapons_assistance#17 / 54↑78.3Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#80 / 104↑9.521Source ↗official
DelusionEvaldiscourages_harm_prevalence_pct#14 / 16↓52.2Source ↗official
DelusionEvalfacilitates_harm_prevalence_pct#15 / 16↑0Source ↗official
SM-Benchadversarial#1 / 92↓68.29Source ↗official
SM-Bencheq_boundaries#65 / 92↑58.71Source ↗official
SM-Benchoverfit#92 / 92↑4.37Source ↗official
SpeechMap model completioncomplete_pct#167 / 181↑32.2Source ↗official
VETO Misfired Alignmentmisfired_alignment_rate_pct#25 / 25↓18.9Source ↗official