← Models

Model profile

Nemotron 3 Super 120B A12B

NVIDIAdeveloper
2026-03-10release date
#50 / 267overall rank
7eval lineages
2discovery sources

Evidence summary

Nemotron 3 Super 120B A12B has an estimated overall rank of #50; its 90% source-sensitivity interval is #20–#179. Its behavior-only rank is #53; company governance moves the combined estimate to #50. Published evidence spans 7 evals and 5 of 7 behavior components. Its strongest relative result is RefusalBench (youden_j, #3 of 19); its weakest is SM-Bench (overfit, #70 of 73).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#219 / 3110.8701↓ lowerSource ↗official
BullshitBench v2clear_pushback_rate#32 / 1050.4375↑ higherSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#57 / 10540.43↓ lowerSource ↗official
DystopiaBenchbasaglia_score#29 / 5065.33↓ lowerSource ↗official
DystopiaBenchbaudrillard_score#22 / 5053.23↓ lowerSource ↗official
DystopiaBenchhuxley_score#23 / 5069.47↓ lowerSource ↗official
DystopiaBenchlaguardia_score#26 / 5067.17↓ lowerSource ↗official
DystopiaBenchorwell_score#25 / 5068.63↓ lowerSource ↗official
DystopiaBenchpetrov_score#34 / 5075.17↓ lowerSource ↗official
RefusalBenchyouden_j#3 / 190.8891↑ higherSource ↗official
SM-Benchadversarial#68 / 7371.22↑ higherSource ↗official
SM-Benchambiguous_interpretation#69 / 7363.39↑ higherSource ↗official
SM-Benchanti_hallucination#42 / 7390.58↑ higherSource ↗official
SM-Bencheq_boundaries#64 / 7352.53↑ higherSource ↗official
SM-Benchoverfit#70 / 7321.86↑ higherSource ↗official
TACbase_welfare_rate#27 / 6829.5↑ higherSource ↗self run