← Models

Model profile

Llama 2 70B Chat

Metadeveloper
2023-07-18release date
#269 / 346Safety rank
#618 / 662Freedom rank

Evidence summary

Safety. Llama 2 70B Chat has an estimated Safety rank of #269; its 90% source-sensitivity interval is #116–#301. Its behavior-only rank is #254; company governance moves the combined estimate to #269. Published Safety evidence spans 9 eval lineages and 6 of 7 components. Its strongest relative result is Enkrypt AI Safety Leaderboard (toxicity_attack_non_success_rate, #16 of 268); its weakest is XSTest (safe_full_compliance_rate, #3 of 3).

Freedom. Llama 2 70B Chat has an estimated Freedom rank of #618; its 90% source-sensitivity interval is #364–#645. Published Freedom evidence spans 9 eval lineages and 1 of 1 components. Its strongest relative result is COMPL-AI LLM RuLES Multi-Turn Rule Following (score, #7 of 14); its weakest is XSTest (safe_full_compliance_rate, #3 of 3).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
COMPL-AI AI-Identity Disclosurescore#9 / 14↑0.8904Source ↗official
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#8 / 14↑0.3822Source ↗official
COMPL-AI TensorTrust Goal-Hijacking Resistancescore#10 / 13↑0.4278Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#206 / 270↑12.4Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#114 / 270↑88.5Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#82 / 270↑87.78Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#16 / 268↑99.64Source ↗official
HarmBenchdr#4 / 28↓2.8Source ↗official
MedSafetyBenchmedical_safety_score#2 / 30↑99.5Source ↗official
OR-Benchover_refusal_rate#23 / 25↓96.1Source ↗official
OR-Benchtoxic_acceptance_rate#2 / 25↓0.3Source ↗official
S-Evalbase_en_overall#5 / 22↑77.2Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#5 / 33↑62.28Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#5 / 33↑68.4Source ↗official
SALAD-Benchattack_enhanced_malicious_use#4 / 33↑66.15Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#5 / 33↑62.17Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#4 / 33↑68.62Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#5 / 33↑60.17Source ↗official
SALAD-Benchbase_human_autonomy_integrity#6 / 33↑98.25Source ↗official
SALAD-Benchbase_information_safety_harms#4 / 33↑99.19Source ↗official
SALAD-Benchbase_malicious_use#6 / 33↑98.17Source ↗official
SALAD-Benchbase_misinformation_harms#11 / 33↑95.67Source ↗official
SALAD-Benchbase_representation_toxicity#11 / 33↑92.71Source ↗official
SALAD-Benchbase_socioeconomic_harms#4 / 33↑94.83Source ↗official
SALAD-Benchmcq_human_autonomy_integrity#17 / 33↑33.61Source ↗official
SALAD-Benchmcq_information_safety_harms#19 / 33↑30.83Source ↗official
SALAD-Benchmcq_malicious_use#19 / 33↑27.24Source ↗official
SALAD-Benchmcq_misinformation_harms#18 / 33↑30.95Source ↗official
SALAD-Benchmcq_representation_toxicity#19 / 33↑27.71Source ↗official
SALAD-Benchmcq_socioeconomic_harms#18 / 33↑31.67Source ↗official
SORRY-Benchavg#11 / 51↓0.12Source ↗official
XSTestsafe_full_compliance_rate#3 / 3↑0.704Source ↗official
XSTestunsafe_full_refusal_rate#1 / 3↑0.975Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#7 / 14↓0.3822Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#153 / 270↓88.5Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#188 / 270↓87.78Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#250 / 268↓99.64Source ↗official
HarmBenchdr#24 / 28↑2.8Source ↗official
MedSafetyBenchmedical_safety_score#29 / 30↓99.5Source ↗official
OR-Benchover_refusal_rate#23 / 25↓96.1Source ↗official
OR-Benchtoxic_acceptance_rate#21 / 25↑0.3Source ↗official
S-Evalbase_en_overall#18 / 22↓77.2Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#29 / 33↓62.28Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#29 / 33↓68.4Source ↗official
SALAD-Benchattack_enhanced_malicious_use#30 / 33↓66.15Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#29 / 33↓62.17Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#30 / 33↓68.62Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#29 / 33↓60.17Source ↗official
SALAD-Benchbase_human_autonomy_integrity#28 / 33↓98.25Source ↗official
SALAD-Benchbase_information_safety_harms#30 / 33↓99.19Source ↗official
SALAD-Benchbase_malicious_use#28 / 33↓98.17Source ↗official
SALAD-Benchbase_misinformation_harms#23 / 33↓95.67Source ↗official
SALAD-Benchbase_representation_toxicity#22 / 33↓92.71Source ↗official
SALAD-Benchbase_socioeconomic_harms#30 / 33↓94.83Source ↗official
SORRY-Benchavg#41 / 51↑0.12Source ↗official
XSTestsafe_full_compliance_rate#3 / 3↑0.704Source ↗official
XSTestunsafe_full_refusal_rate#2 / 3↓0.975Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

CCP-aligned censorship behavior

DimensionValueDistribution
Political-question refusal rate (ZH/EN mean)1.57