← Models

Model profile

Llama 3 70B Instruct

Metadeveloper
2024-04-18release date
#239 / 346Safety rank
#197 / 662Freedom rank

Evidence summary

Safety. Llama 3 70B Instruct has an estimated Safety rank of #239; its 90% source-sensitivity interval is #130–#271. Its behavior-only rank is #224; company governance moves the combined estimate to #239. Published Safety evidence spans 13 eval lineages and 7 of 7 components. Its strongest relative result is Large-scale Moral Machine experiment on LLMs (human_choice_distance, #5 of 39); its weakest is AgentDojo (utility_under_attack, #15 of 15).

Freedom. Llama 3 70B Instruct has an estimated Freedom rank of #197; its 90% source-sensitivity interval is #164–#302. Published Freedom evidence spans 10 eval lineages and 1 of 1 components. Its strongest relative result is OR-Bench (toxic_acceptance_rate, #3 of 25); its weakest is Enkrypt AI Safety Leaderboard (toxicity_attack_non_success_rate, #227 of 268).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#256 / 358↓0.8754Source ↗official
AgentDojotargeted_attack_success_rate#11 / 15↓0.256Source ↗official
AgentDojoutility_under_attack#15 / 15↑0.1828Source ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#45 / 80↑0.646Source ↗official
CASE-Benchagreement_accuracy#3 / 7↑84.44Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#135 / 270↑16.02Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#127 / 270↑87.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#82 / 270↑87.78Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#42 / 268↑98.95Source ↗official
HELM Safetyanthropic_red_team#66 / 80↑0.967Source ↗official
HELM Safetybbq#53 / 80↑0.91Source ↗official
HELM Safetyharmbench#57 / 80↑0.64Source ↗official
HELM Safetysimple_safety_tests#32 / 80↑0.99Source ↗official
HELM Safetyxstest#29 / 80↑0.968Source ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#5 / 39↓0.7475Source ↗official
MT-JailBench CrescendoXsafety_score#17 / 21↑11.32Source ↗official
OR-Benchover_refusal_rate#9 / 25↓37.7Source ↗official
OR-Benchtoxic_acceptance_rate#22 / 25↓21.3Source ↗official
S-Evalbase_en_overall#15 / 22↑54.7Source ↗official
SORRY-Benchavg#31 / 51↓0.35Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AIRBench 2024 Safety Scenariossafety_scenarios#36 / 80↓0.646Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#142 / 270↓87.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#188 / 270↓87.78Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#227 / 268↓98.95Source ↗official
HELM Safetyanthropic_red_team#15 / 80↓0.967Source ↗official
HELM Safetyharmbench#24 / 80↓0.64Source ↗official
HELM Safetysimple_safety_tests#42 / 80↓0.99Source ↗official
HELM Safetyxstest#29 / 80↑0.968Source ↗official
MT-JailBench CrescendoXsafety_score#3 / 21↓11.32Source ↗official
OR-Benchover_refusal_rate#9 / 25↓37.7Source ↗official
OR-Benchtoxic_acceptance_rate#3 / 25↑21.3Source ↗official
S-Evalbase_en_overall#8 / 22↓54.7Source ↗official
SORRY-Benchavg#19 / 51↑0.35Source ↗official
SpeechMap model completioncomplete_pct#55 / 181↑70.8Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism64.7
Self-direction47.2
Care / Harm12.3
Fairness / Cheating11.9
Ethical88.2