← Models

Model profile

Llama 3 70B Instruct

Metadeveloper
2024-04-18release date
#185 / 267overall rank
13eval lineages

Evidence summary

Llama 3 70B Instruct has an estimated overall rank of #185; its 90% source-sensitivity interval is #93–#220. Its behavior-only rank is #171; company governance moves the combined estimate to #185. Published evidence spans 13 evals and 7 of 7 behavior components. Its strongest relative result is Enkrypt AI Safety Leaderboard (toxicity_attack_non_success_rate, #1 of 258); its weakest is AgentDojo (utility_under_attack, #15 of 15).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#209 / 3110.865↓ lowerSource ↗official
AgentDojotargeted_attack_success_rate#11 / 150.256↓ lowerSource ↗official
AgentDojoutility_under_attack#15 / 150.1828↑ higherSource ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#45 / 800.646↑ higherSource ↗official
CASE-Benchagreement_accuracy#3 / 784.44↑ higherSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#92 / 26019.64↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#168 / 26086.33↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#63 / 26087.78↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#1 / 258100↑ higherSource ↗official
HELM Safetyanthropic_red_team#66 / 800.967↑ higherSource ↗official
HELM Safetybbq#53 / 800.91↑ higherSource ↗official
HELM Safetyharmbench#57 / 800.64↑ higherSource ↗official
HELM Safetysimple_safety_tests#32 / 800.99↑ higherSource ↗official
HELM Safetyxstest#29 / 800.968↑ higherSource ↗official
Large-scale Moral Machine experiment on LLMshuman_choice_distance#5 / 390.7475↓ lowerSource ↗official
OR-Benchover_refusal_rate#9 / 2537.7↓ lowerSource ↗official
OR-Benchtoxic_acceptance_rate#22 / 2521.3↓ lowerSource ↗official
S-Evalbase_en_overall#15 / 2254.7↑ higherSource ↗official
SORRY-Benchavg#31 / 510.35↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

ValueCompass

DimensionValueDistribution
Universalism64.7
Self-direction47.2
Care / Harm12.3
Fairness / Cheating11.9
Ethical88.2