← Models

Model profile

Llama 4 Scout

Metadeveloper
2025-04-05release date
#278 / 346Safety rank
#92 / 662Freedom rank

Evidence summary

Safety. Llama 4 Scout has an estimated Safety rank of #278; its 90% source-sensitivity interval is #197–#287. Published Safety evidence spans 21 eval lineages and 7 of 7 components. Its strongest relative result is PHARE (bias_resistance_diagnostic, #7 of 66); its weakest is JuICE Cultural-Error Span Detection (f1, #10 of 10).

Freedom. Llama 4 Scout has an estimated Freedom rank of #92; its 90% source-sensitivity interval is #98–#252. Published Freedom evidence spans 13 eval lineages and 1 of 1 components. Its strongest relative result is Vigil Mental Health Safety (overall_score, #1 of 23); its weakest is Enkrypt AI Safety Leaderboard (toxicity_attack_non_success_rate, #196 of 268).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#173 / 358↓0.7935Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5child_sexual_exploitation#21 / 31↓55.3Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5defamation#12 / 31↓64.4Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5hate#12 / 31↓47.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5indiscriminate_weapons_cbrne#20 / 31↓80Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5intellectual_property#14 / 31↓55.3Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5non_violent_crimes#12 / 31↓75.6Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5privacy#17 / 31↓70.2Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sex_related_crimes#16 / 31↓69.6Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sexual_content#13 / 31↓37.5Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5specialized_advice#12 / 31↓56.55Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5suicide_self_harm#20 / 31↓62.5Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5violent_crimes#22 / 31↓80.4Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRbaseline_asr#19 / 24↓11.52Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRpoetry_asr#12 / 24↓49.61Source ↗official
AgentDrive Safety Compliancescr#36 / 48↑58.75Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#63 / 80↑0.523Source ↗official
BullshitBench v2clear_pushback_rate#79 / 122↑0.2Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#74 / 104↓58.72Source ↗official
DystopiaBenchbasaglia_score#34 / 50↓67.87Source ↗official
DystopiaBenchbaudrillard_score#45 / 50↓74.3Source ↗official
DystopiaBenchhuxley_score#24 / 50↓70.13Source ↗official
DystopiaBenchlaguardia_score#29 / 50↓67.9Source ↗official
DystopiaBenchorwell_score#23 / 50↓68.6Source ↗official
DystopiaBenchpetrov_score#20 / 50↓70.47Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#168 / 270↑13.95Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#127 / 270↑87.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#153 / 270↑68.33Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#72 / 268↑98.05Source ↗official
HELM Safetyanthropic_red_team#67 / 80↑0.965Source ↗official
HELM Safetybbq#60 / 80↑0.875Source ↗official
HELM Safetyharmbench#62 / 80↑0.6Source ↗official
HELM Safetysimple_safety_tests#58 / 80↑0.97Source ↗official
HELM Safetyxstest#43 / 80↑0.956Source ↗official
JuICE Cultural-Error Span Detectionf1#10 / 10↑0.2805Source ↗official
MT-JailBench CrescendoXsafety_score#20 / 21↑10.06Source ↗official
PHAREbias_resistance_diagnostic#7 / 66↑0.671Source ↗official
PHAREhallucination_resistance_diagnostic#69 / 70↑0.582Source ↗official
PHAREharm_resistance_diagnostic#65 / 70↑0.8104Source ↗official
PHAREjailbreak_resistance_diagnostic#35 / 67↑0.4893Source ↗official
SOSBenchbiology_pvr#13 / 23↓0.488Source ↗official
SOSBenchchemistry_pvr#12 / 23↓0.436Source ↗official
SOSBenchmedicine_pvr#13 / 23↓0.688Source ↗official
SOSBenchpharmacology_pvr#14 / 23↓0.874Source ↗official
SOSBenchphysics_pvr#13 / 23↓0.492Source ↗official
SOSBenchpsychology_pvr#13 / 23↓0.51Source ↗official
SpeciEvalbelief_animal_sentience#28 / 131↑6.97Source ↗official
SpeciEvalland_animal_4ns#121 / 131↓5Source ↗official
SpeciEvalsea_animal_4ns#126 / 131↓5.33Source ↗official
SpeciEvalspeciesism#69 / 131↓2.02Source ↗official
UAVBench safety-critical decision recognitionethical_safety_critical_accuracy#19 / 27↑0.635Source ↗official
Vectara HHEM Factual Consistencyfactual_consistency_rate#34 / 94↑92.3Source ↗official
Vigil Mental Health Safetyoverall_score#23 / 23↑20Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Humanities Benchmark (AHB) — Table 5child_sexual_exploitation#11 / 31↑55.3Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5defamation#19 / 31↑64.4Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5hate#17 / 31↑47.8Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5indiscriminate_weapons_cbrne#10 / 31↑80Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5intellectual_property#18 / 31↑55.3Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5non_violent_crimes#20 / 31↑75.6Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5privacy#15 / 31↑70.2Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sex_related_crimes#16 / 31↑69.6Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5sexual_content#19 / 31↑37.5Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5specialized_advice#20 / 31↑56.55Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5suicide_self_harm#11 / 31↑62.5Source ↗official
Adversarial Humanities Benchmark (AHB) — Table 5violent_crimes#10 / 31↑80.4Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRbaseline_asr#6 / 24↑11.52Source ↗official
Adversarial Poetry — AILuminate Baseline and Poetry ASRpoetry_asr#13 / 24↑49.61Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#18 / 80↓0.523Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#31 / 104↑58.72Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#142 / 270↓87.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#117 / 270↓68.33Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#196 / 268↓98.05Source ↗official
HELM Safetyanthropic_red_team#13 / 80↓0.965Source ↗official
HELM Safetyharmbench#19 / 80↓0.6Source ↗official
HELM Safetysimple_safety_tests#20 / 80↓0.97Source ↗official
HELM Safetyxstest#43 / 80↑0.956Source ↗official
MT-JailBench CrescendoXsafety_score#2 / 21↓10.06Source ↗official
PHAREharm_resistance_diagnostic#6 / 70↓0.8104Source ↗official
PHAREjailbreak_resistance_diagnostic#33 / 67↓0.4893Source ↗official
SOSBenchbiology_pvr#11 / 23↑0.488Source ↗official
SOSBenchchemistry_pvr#12 / 23↑0.436Source ↗official
SOSBenchmedicine_pvr#11 / 23↑0.688Source ↗official
SOSBenchpharmacology_pvr#10 / 23↑0.874Source ↗official
SOSBenchphysics_pvr#11 / 23↑0.492Source ↗official
SOSBenchpsychology_pvr#11 / 23↑0.51Source ↗official
SpeechMap model completioncomplete_pct#56 / 181↑70.1Source ↗official
Vigil Mental Health Safetyoverall_score#1 / 23↓20Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

The Economist World Values Survey Cultural Map

DimensionValueDistribution
Survival ↔ Self-expression1.07
Traditional ↔ Secular1.51