← Models

Model profile

Llama 3.2 1B Instruct

Metadeveloper
2024-09-25release date
#253 / 267overall rank
7eval lineages

Evidence summary

Llama 3.2 1B Instruct has an estimated overall rank of #253; its 90% source-sensitivity interval is #191–#266. Its behavior-only rank is #244; company governance moves the combined estimate to #253. Published evidence spans 7 evals and 6 of 7 behavior components. Its strongest relative result is Enkrypt AI Safety Leaderboard (toxicity_attack_non_success_rate, #53 of 258); its weakest is Enkrypt AI Safety Leaderboard (bias_attack_non_success_rate, #250 of 260).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#275 / 3110.9155↓ lowerSource ↗official
BlueBench AttaQ-100attaq_harmlessness_reward_pct#14 / 1883↑ higherSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#38 / 10521.39↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#250 / 2604.13↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#131 / 26088.33↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#54 / 26091.11↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#53 / 25898.82↑ higherSource ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#25 / 391.182↓ lowerSource ↗official
Open LLM Safety Indexjailbreakbench_safety_rate#14 / 210.2667↑ higherSource ↗official
Open LLM Safety Indexstrongreject_safety_rate#14 / 210.1333↑ higherSource ↗official
PandaBench JBB direct-request panelsafety_rate#14 / 460.99↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-6
Government56.9
Diplomacy56.1
Economy53
Society45.4