← Models

Model profile

Nemotron 3 Super 120B A12B

NVIDIAdeveloper
2026-03-10release date
#196 / 346Safety rank
#533 / 662Freedom rank
2discovery sources

Evidence summary

Safety. Nemotron 3 Super 120B A12B has an estimated Safety rank of #196; its 90% source-sensitivity interval is #107–#279. Its behavior-only rank is #197; company governance moves the combined estimate to #196. Published Safety evidence spans 11 eval lineages and 6 of 7 components. Its strongest relative result is Enkrypt AI Safety Leaderboard (harmful_attack_non_success_rate, #63 of 270); its weakest is Pokee-Isaac model card — DTAP (benign_task_success_rate, #6 of 6).

Freedom. Nemotron 3 Super 120B A12B has an estimated Freedom rank of #533; its 90% source-sensitivity interval is #336–#634. Published Freedom evidence spans 5 eval lineages and 1 of 1 components. Its strongest relative result is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #13 of 270); its weakest is SM-Bench (overfit, #89 of 92).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#248 / 358↓0.8696Source ↗official
Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating#99 / 111↑1397.0Source ↗official
BullshitBench v2clear_pushback_rate#46 / 122↑0.4375Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#56 / 104↓40.43Source ↗official
DystopiaBenchbasaglia_score#29 / 50↓65.33Source ↗official
DystopiaBenchbaudrillard_score#22 / 50↓53.23Source ↗official
DystopiaBenchhuxley_score#23 / 50↓69.47Source ↗official
DystopiaBenchlaguardia_score#26 / 50↓67.17Source ↗official
DystopiaBenchorwell_score#25 / 50↓68.63Source ↗official
DystopiaBenchpetrov_score#34 / 50↓75.17Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#158 / 270↑14.73Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#257 / 270↑50.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#63 / 270↑92.22Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#63 / 268↑98.18Source ↗official
MT-JailBench CrescendoXsafety_score#5 / 21↑44.03Source ↗official
Pokee-Isaac model card — DTAPbenign_task_success_rate#6 / 6↑0.633Source ↗official
Pokee-Isaac model card — DTAPcombined_attack_success_rate#5 / 6↓0.604Source ↗official
RefusalBenchyouden_j#12 / 19↑0.06383Source ↗official
SM-Benchadversarial#87 / 92↑71.22Source ↗official
SM-Benchambiguous_interpretation#88 / 92↑63.39Source ↗official
SM-Benchanti_hallucination#60 / 92↑90.58Source ↗official
SM-Bencheq_boundaries#83 / 92↑52.53Source ↗official
SM-Benchoverfit#89 / 92↑21.86Source ↗official
TACbase_welfare_rate#35 / 92↑29.5Source ↗self run

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#49 / 104↑40.43Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#13 / 270↓50.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#206 / 270↓92.22Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#201 / 268↓98.18Source ↗official
MT-JailBench CrescendoXsafety_score#17 / 21↓44.03Source ↗official
SM-Benchadversarial#6 / 92↓71.22Source ↗official
SM-Bencheq_boundaries#83 / 92↑52.53Source ↗official
SM-Benchoverfit#89 / 92↑21.86Source ↗official
SpeechMap model completioncomplete_pct#136 / 181↑42.5Source ↗official