← Models

Model profile

Llama 2 70B Chat

Metadeveloper
2023-07-18release date
#192 / 267overall rank
7eval lineages

Evidence summary

Llama 2 70B Chat has an estimated overall rank of #192; its 90% source-sensitivity interval is #63–#233. Its behavior-only rank is #182; company governance moves the combined estimate to #192. Published evidence spans 7 evals and 6 of 7 behavior components. Its strongest relative result is OR-Bench (toxic_acceptance_rate, #2 of 25); its weakest is XSTest (safe_full_compliance_rate, #3 of 3).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#195 / 26012.4↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#125 / 26088.5↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#63 / 26087.78↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#21 / 25899.64↑ higherSource ↗official
HarmBenchdr#4 / 282.8↓ lowerSource ↗official
OR-Benchover_refusal_rate#23 / 2596.1↓ lowerSource ↗official
OR-Benchtoxic_acceptance_rate#2 / 250.3↓ lowerSource ↗official
S-Evalbase_en_overall#5 / 2277.2↑ higherSource ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#5 / 3362.28↑ higherSource ↗official
SALAD-Benchattack_enhanced_information_safety_harms#5 / 3368.4↑ higherSource ↗official
SALAD-Benchattack_enhanced_malicious_use#4 / 3366.15↑ higherSource ↗official
SALAD-Benchattack_enhanced_misinformation_harms#5 / 3362.17↑ higherSource ↗official
SALAD-Benchattack_enhanced_representation_toxicity#4 / 3368.62↑ higherSource ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#5 / 3360.17↑ higherSource ↗official
SALAD-Benchbase_human_autonomy_integrity#6 / 3398.25↑ higherSource ↗official
SALAD-Benchbase_information_safety_harms#4 / 3399.19↑ higherSource ↗official
SALAD-Benchbase_malicious_use#6 / 3398.17↑ higherSource ↗official
SALAD-Benchbase_misinformation_harms#11 / 3395.67↑ higherSource ↗official
SALAD-Benchbase_representation_toxicity#11 / 3392.71↑ higherSource ↗official
SALAD-Benchbase_socioeconomic_harms#4 / 3394.83↑ higherSource ↗official
SALAD-Benchmcq_human_autonomy_integrity#17 / 3333.61↑ higherSource ↗official
SALAD-Benchmcq_information_safety_harms#19 / 3330.83↑ higherSource ↗official
SALAD-Benchmcq_malicious_use#19 / 3327.24↑ higherSource ↗official
SALAD-Benchmcq_misinformation_harms#18 / 3330.95↑ higherSource ↗official
SALAD-Benchmcq_representation_toxicity#19 / 3327.71↑ higherSource ↗official
SALAD-Benchmcq_socioeconomic_harms#18 / 3331.67↑ higherSource ↗official
SORRY-Benchavg#11 / 510.12↓ lowerSource ↗official
XSTestsafe_full_compliance_rate#3 / 30.704↑ higherSource ↗official
XSTestunsafe_full_refusal_rate#1 / 30.975↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

CCP-aligned censorship behavior

DimensionValueDistribution
Political-question refusal rate (ZH/EN mean)1.57