← Models

Model profile

Llama 3.1 405B Instruct

Metadeveloper
2024-07-23release date
#212 / 346Safety rank
#379 / 662Freedom rank

Evidence summary

Safety. Llama 3.1 405B Instruct has an estimated Safety rank of #212; its 90% source-sensitivity interval is #135–#272. Its behavior-only rank is #201; company governance moves the combined estimate to #212. Published Safety evidence spans 23 eval lineages and 6 of 7 components. Its strongest relative result is PHARE (bias_resistance_diagnostic, #3 of 66); its weakest is BlueBench AttaQ-100 (attaq_harmlessness_reward_pct, #17 of 18).

Freedom. Llama 3.1 405B Instruct has an estimated Freedom rank of #379; its 90% source-sensitivity interval is #190–#510. Published Freedom evidence spans 18 eval lineages and 1 of 1 components. Its strongest relative result is PandaBench JBB direct-request panel (safety_rate, #4 of 46); its weakest is PHARE (jailbreak_resistance_diagnostic, #63 of 67).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#93 / 358↓0.5244Source ↗official
AbstentionBenchanswer_unknown_f1#11 / 20↑0.87Source ↗official
AbstentionBenchfalse_premise_f1#7 / 20↑0.7164Source ↗official
AbstentionBenchstale_f1#9 / 20↑0.6419Source ↗official
AbstentionBenchsubjective_f1#10 / 20↑0.7422Source ↗official
AbstentionBenchunderspecified_context_f1#3 / 20↑0.7157Source ↗official
AbstentionBenchunderspecified_intent_f1#14 / 20↑0.7024Source ↗official
Adversarial Robustnessscore#2 / 8↓10Source ↗official
Agent-SafetyBenchcompromise_availability#14 / 16↑19.6Source ↗official
Agent-SafetyBenchharmful_vulnerable_code#6 / 16↑40.4Source ↗official
Agent-SafetyBenchleak_sensitive_information#12 / 16↑25.2Source ↗official
Agent-SafetyBenchphysical_harm#12 / 16↑24.4Source ↗official
Agent-SafetyBenchproduce_unsafe_information#4 / 16↑98.8Source ↗official
Agent-SafetyBenchproperty_loss#14 / 16↑27.6Source ↗official
Agent-SafetyBenchspread_unsafe_information#6 / 16↑14.4Source ↗official
Agent-SafetyBenchviolate_law_ethics#5 / 16↑32.8Source ↗official
AgentHarmharm_score#2 / 12↓4.3Source ↗official
AILuminate General Purpose AI Chatchild_sexual_exploitation#16 / 32↓20.4Source ↗official
AILuminate General Purpose AI Chatdefamation#21 / 32↓20.6Source ↗official
AILuminate General Purpose AI Chathate#24 / 32↓14.5Source ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#21 / 32↓29.5Source ↗official
AILuminate General Purpose AI Chatintellectual_property#23 / 32↓19Source ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#23 / 32↓23.9Source ↗official
AILuminate General Purpose AI Chatprivacy#20 / 32↓17.9Source ↗official
AILuminate General Purpose AI Chatsex_related_crimes#22 / 32↓19.9Source ↗official
AILuminate General Purpose AI Chatsexual_content#24 / 32↓35.2Source ↗official
AILuminate General Purpose AI Chatspecialized_advice#21 / 32↓20.2Source ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#21 / 32↓15.7Source ↗official
AILuminate General Purpose AI Chatviolent_crimes#23 / 32↓30.9Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#52 / 80↑0.586Source ↗official
Alignment Leaderboardcorrigibility#20 / 24↑4.058Source ↗official
Alignment Leaderboardhonesty#17 / 24↑3.475Source ↗official
Alignment Leaderboardnon_manipulation#20 / 24↑3.034Source ↗official
Alignment Leaderboardrobustness#15 / 24↑3.36Source ↗official
Alignment Leaderboardsafety#17 / 24↑3.692Source ↗official
Alignment Leaderboardscheming#17 / 24↑3.53Source ↗official
BlueBench AttaQ-100attaq_harmlessness_reward_pct#17 / 18↑81.35Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#44 / 104↓31.56Source ↗official
Confabulationsconfabulation_rate#14 / 52↓14.36Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#189 / 270↑12.92Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#83 / 270↑90.17Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#138 / 270↑73.33Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#145 / 268↑96.05Source ↗official
FORTRESSaverage_risk_score#25 / 60↓20.61Source ↗official
FORTRESSover_refusal_score#38 / 59↓6.03Source ↗official
HELM Safetyanthropic_red_team#67 / 80↑0.965Source ↗official
HELM Safetybbq#33 / 80↑0.945Source ↗official
HELM Safetyharmbench#59 / 80↑0.627Source ↗official
HELM Safetysimple_safety_tests#40 / 80↑0.988Source ↗official
HELM Safetyxstest#41 / 80↑0.959Source ↗official
MASKlying_probability_pct#23 / 53↓33.3Source ↗official
PandaBench JBB direct-request panelsafety_rate#43 / 46↑0.92Source ↗official
PHAREbias_resistance_diagnostic#3 / 66↑0.7523Source ↗official
PHAREhallucination_resistance_diagnostic#55 / 70↑0.674Source ↗official
PHAREharm_resistance_diagnostic#57 / 70↑0.8649Source ↗official
PHAREjailbreak_resistance_diagnostic#5 / 67↑0.7655Source ↗official
SORRY-Benchavg#20 / 51↓0.24Source ↗official
SOSBenchbiology_pvr#14 / 23↓0.59Source ↗official
SOSBenchchemistry_pvr#13 / 23↓0.468Source ↗official
SOSBenchmedicine_pvr#14 / 23↓0.69Source ↗official
SOSBenchpharmacology_pvr#11 / 23↓0.764Source ↗official
SOSBenchphysics_pvr#11 / 23↓0.444Source ↗official
SOSBenchpsychology_pvr#14 / 23↓0.568Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Robustnessscore#7 / 8↑10Source ↗official
Agent-SafetyBenchproduce_unsafe_information#13 / 16↓98.8Source ↗official
AILuminate General Purpose AI Chatchild_sexual_exploitation#17 / 32↑20.4Source ↗official
AILuminate General Purpose AI Chatdefamation#12 / 32↑20.6Source ↗official
AILuminate General Purpose AI Chathate#9 / 32↑14.5Source ↗official
AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne#12 / 32↑29.5Source ↗official
AILuminate General Purpose AI Chatintellectual_property#10 / 32↑19Source ↗official
AILuminate General Purpose AI Chatnon_violent_crimes#10 / 32↑23.9Source ↗official
AILuminate General Purpose AI Chatprivacy#12 / 32↑17.9Source ↗official
AILuminate General Purpose AI Chatsex_related_crimes#11 / 32↑19.9Source ↗official
AILuminate General Purpose AI Chatsexual_content#9 / 32↑35.2Source ↗official
AILuminate General Purpose AI Chatspecialized_advice#12 / 32↑20.2Source ↗official
AILuminate General Purpose AI Chatsuicide_self_harm#12 / 32↑15.7Source ↗official
AILuminate General Purpose AI Chatviolent_crimes#10 / 32↑30.9Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#29 / 80↓0.586Source ↗official
BlueBench AttaQ-100attaq_harmlessness_reward_pct#2 / 18↓81.35Source ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#61 / 104↑31.56Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#185 / 270↓90.17Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#130 / 270↓73.33Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#122 / 268↓96.05Source ↗official
FORTRESSaverage_risk_score#36 / 60↑20.61Source ↗official
FORTRESSover_refusal_score#38 / 59↓6.03Source ↗official
HELM Safetyanthropic_red_team#13 / 80↓0.965Source ↗official
HELM Safetyharmbench#22 / 80↓0.627Source ↗official
HELM Safetysimple_safety_tests#39 / 80↓0.988Source ↗official
HELM Safetyxstest#41 / 80↑0.959Source ↗official
PandaBench JBB direct-request panelsafety_rate#4 / 46↓0.92Source ↗official
PHAREharm_resistance_diagnostic#14 / 70↓0.8649Source ↗official
PHAREjailbreak_resistance_diagnostic#63 / 67↓0.7655Source ↗official
SORRY-Benchavg#31 / 51↑0.24Source ↗official
SOSBenchbiology_pvr#10 / 23↑0.59Source ↗official
SOSBenchchemistry_pvr#11 / 23↑0.468Source ↗official
SOSBenchmedicine_pvr#10 / 23↑0.69Source ↗official
SOSBenchpharmacology_pvr#13 / 23↑0.764Source ↗official
SOSBenchphysics_pvr#13 / 23↑0.444Source ↗official
SOSBenchpsychology_pvr#10 / 23↑0.568Source ↗official
SpeechMap model completioncomplete_pct#104 / 181↑51.7Source ↗official
UGI Leaderboard — base-model willingnesswillingness_adherence_score#144 / 156↑0.5Source ↗official
UGI Leaderboard — base-model willingnesswillingness_direct_score#113 / 156↑2Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-24.2
Government45.1
Diplomacy69.2
Economy40.8
Society63.8

ValueCompass

DimensionValueDistribution
Universalism66.2
Self-direction48.4
Care / Harm18.5
Fairness / Cheating17.4
Ethical88.6