← Models

Model profile

Llama 2 13B Chat

Metadeveloper
2023-07-18release date
#220 / 346Safety rank
#602 / 662Freedom rank

Evidence summary

Safety. Llama 2 13B Chat has an estimated Safety rank of #220; its 90% source-sensitivity interval is #124–#276. Its behavior-only rank is #205; company governance moves the combined estimate to #220. Published Safety evidence spans 10 eval lineages and 6 of 7 components. Its strongest relative result is COMPL-AI AI-Identity Disclosure (score, #1 of 14); its weakest is SafetyBench (OFF, #20 of 21).

Freedom. Llama 2 13B Chat has an estimated Freedom rank of #602; its 90% source-sensitivity interval is #316–#639. Published Freedom evidence spans 9 eval lineages and 1 of 1 components. Its strongest relative result is COMPL-AI LLM RuLES Multi-Turn Rule Following (score, #6 of 14); its weakest is S-Eval (base_en_overall, #21 of 22).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
COMPL-AI AI-Identity Disclosurescore#1 / 14↑1Source ↗official
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#9 / 14↑0.3652Source ↗official
COMPL-AI TensorTrust Goal-Hijacking Resistancescore#11 / 13↑0.4175Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#223 / 270↑11.11Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#51 / 270↑91.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#80 / 270↑88.33Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#20 / 268↑99.55Source ↗official
HarmBenchdr#4 / 28↓2.8Source ↗official
JailBenchjailbreak_success_rate#6 / 14↓55.39Source ↗official
MedSafetyBenchmedical_safety_score#3 / 30↑99.25Source ↗official
OR-Benchover_refusal_rate#20 / 25↓91Source ↗official
OR-Benchtoxic_acceptance_rate#2 / 25↓0.3Source ↗official
S-Evalbase_en_overall#2 / 22↑85.1Source ↗official
SafetyBenchEM#17 / 21↑54.6Source ↗official
SafetyBenchIA#15 / 21↑68.5Source ↗official
SafetyBenchMH#15 / 21↑73.6Source ↗official
SafetyBenchOFF#20 / 21↑48.4Source ↗official
SafetyBenchPH#16 / 21↑60.7Source ↗official
SafetyBenchPP#16 / 21↑70.1Source ↗official
SafetyBenchUB#5 / 21↑66.3Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#4 / 33↑62.72Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#4 / 33↑71.01Source ↗official
SALAD-Benchattack_enhanced_malicious_use#5 / 33↑62.56Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#4 / 33↑70.23Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#5 / 33↑65.85Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#4 / 33↑68.4Source ↗official
SALAD-Benchbase_human_autonomy_integrity#4 / 33↑98.43Source ↗official
SALAD-Benchbase_information_safety_harms#16 / 33↑94.92Source ↗official
SALAD-Benchbase_malicious_use#7 / 33↑97.98Source ↗official
SALAD-Benchbase_misinformation_harms#4 / 33↑97.64Source ↗official
SALAD-Benchbase_representation_toxicity#4 / 33↑95.71Source ↗official
SALAD-Benchbase_socioeconomic_harms#13 / 33↑91.19Source ↗official
SALAD-Benchmcq_human_autonomy_integrity#26 / 33↑10.28Source ↗official
SALAD-Benchmcq_information_safety_harms#23 / 33↑17.5Source ↗official
SALAD-Benchmcq_malicious_use#26 / 33↑7.885Source ↗official
SALAD-Benchmcq_misinformation_harms#26 / 33↑10.95Source ↗official
SALAD-Benchmcq_representation_toxicity#26 / 33↑8.229Source ↗official
SALAD-Benchmcq_socioeconomic_harms#26 / 33↑12.78Source ↗official
SORRY-Benchavg#15 / 51↓0.15Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
COMPL-AI LLM RuLES Multi-Turn Rule Followingscore#6 / 14↓0.3652Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#219 / 270↓91.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#190 / 270↓88.33Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#247 / 268↓99.55Source ↗official
HarmBenchdr#24 / 28↑2.8Source ↗official
JailBenchjailbreak_success_rate#9 / 14↑55.39Source ↗official
MedSafetyBenchmedical_safety_score#28 / 30↓99.25Source ↗official
OR-Benchover_refusal_rate#20 / 25↓91Source ↗official
OR-Benchtoxic_acceptance_rate#21 / 25↑0.3Source ↗official
S-Evalbase_en_overall#21 / 22↓85.1Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#30 / 33↓62.72Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#30 / 33↓71.01Source ↗official
SALAD-Benchattack_enhanced_malicious_use#29 / 33↓62.56Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#30 / 33↓70.23Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#29 / 33↓65.85Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#30 / 33↓68.4Source ↗official
SALAD-Benchbase_human_autonomy_integrity#30 / 33↓98.43Source ↗official
SALAD-Benchbase_information_safety_harms#18 / 33↓94.92Source ↗official
SALAD-Benchbase_malicious_use#27 / 33↓97.98Source ↗official
SALAD-Benchbase_misinformation_harms#30 / 33↓97.64Source ↗official
SALAD-Benchbase_representation_toxicity#30 / 33↓95.71Source ↗official
SALAD-Benchbase_socioeconomic_harms#20 / 33↓91.19Source ↗official
SORRY-Benchavg#37 / 51↑0.15Source ↗official