← Models

Model profile

Llama 3.1 Nemotron Ultra 253B v1

NVIDIAdeveloper
2025-04-07release date
Not rankedSafety rank
#451 / 662Freedom rank

Evidence summary

Safety. Llama 3.1 Nemotron Ultra 253B v1 does not meet the evidence gate for a Safety rank. Published Safety evidence spans 2 eval lineages and 4 of 7 components. Its strongest relative result is Concordia — AirBench-SecurityRisks (safety_score, #1 of 63); its weakest is Concordia — CyberSecEval2-PromptInjection (safety_score, #53 of 63).

Freedom. Llama 3.1 Nemotron Ultra 253B v1 has an estimated Freedom rank of #451; its 90% source-sensitivity interval is #312–#558. Published Freedom evidence spans 1 eval lineages and 1 of 1 components. Its strongest relative result is Concordia — SOSBench-Bio (safety_score, #24 of 63); its weakest is Concordia — AirBench-SecurityRisks (safety_score, #58 of 63).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#190 / 358↓0.8123Source ↗official
Concordia — AirBench-Deceptionsafety_score#20 / 63↑90.37Source ↗official
Concordia — AirBench-Manipulationsafety_score#23 / 56↑92Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#1 / 63↑100Source ↗official
Concordia — APEsafety_score#27 / 55↑27.81Source ↗official
Concordia — CyberSecEval2-PromptInjectionsafety_score#53 / 63↑72.51Source ↗official
Concordia — DarkBenchsafety_score#11 / 55↑64.44Source ↗official
Concordia — MASKsafety_score#25 / 62↑65.75Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#9 / 63↑98.99Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#24 / 63↑25.77Source ↗official
Concordia — SOSBench-Biosafety_score#39 / 63↑72.6Source ↗official
Concordia — SOSBench-Chemsafety_score#38 / 63↑83.37Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Concordia — AirBench-Deceptionsafety_score#44 / 63↓90.37Source ↗official
Concordia — AirBench-Manipulationsafety_score#33 / 56↓92Source ↗official
Concordia — AirBench-PoliticalPersuasionsafety_score#51 / 56↓87.14Source ↗official
Concordia — AirBench-SecurityRiskssafety_score#58 / 63↓100Source ↗official
Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score#54 / 63↓98.99Source ↗official
Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score#40 / 63↓25.77Source ↗official
Concordia — SOSBench-Biosafety_score#24 / 63↓72.6Source ↗official
Concordia — SOSBench-Chemsafety_score#26 / 63↓83.37Source ↗official