← Models

Model profile

Llama 3.2 3B Instruct

Metadeveloper
2024-09-25release date
#250 / 267overall rank
8eval lineages

Evidence summary

Llama 3.2 3B Instruct has an estimated overall rank of #250; its 90% source-sensitivity interval is #189–#258. Its behavior-only rank is #243; company governance moves the combined estimate to #250. Published evidence spans 8 evals and 6 of 7 behavior components. Its strongest relative result is AA-Omniscience (hallucination_rate, #66 of 311); its weakest is UAVBench safety-critical decision recognition (ethical_safety_critical_accuracy, #27 of 27).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#66 / 3110.5753↓ lowerSource ↗official
BlueBench AttaQ-100attaq_harmlessness_reward_pct#13 / 1883.37↑ higherSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#58 / 10540.58↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#211 / 26011.11↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#177 / 26085.67↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#80 / 26085↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#154 / 25895.5↑ higherSource ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#34 / 391.519↓ lowerSource ↗official
Open LLM Safety Indexjailbreakbench_safety_rate#14 / 210.2667↑ higherSource ↗official
Open LLM Safety Indexstrongreject_safety_rate#5 / 210.6667↑ higherSource ↗official
PandaBench JBB direct-request panelsafety_rate#14 / 460.99↑ higherSource ↗official
UAVBench safety-critical decision recognitionethical_safety_critical_accuracy#27 / 270.475↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-1.2
Government52.4
Diplomacy55.7
Economy44.7
Society49.4