← Models

Model profile

InternLM2 Chat 20B

InternLMdeveloper
2024-01-17release date
#132 / 346Safety rank
#464 / 662Freedom rank

Evidence summary

Safety. InternLM2 Chat 20B has an estimated Safety rank of #132; its 90% source-sensitivity interval is #42–#285. Its behavior-only rank is #128; company governance moves the combined estimate to #132. Published Safety evidence spans 5 eval lineages and 5 of 7 components. Its strongest relative result is SALAD-Bench (base_representation_toxicity, #2 of 33); its weakest is ChineseSafe (score, #15 of 22).

Freedom. InternLM2 Chat 20B has an estimated Freedom rank of #464; its 90% source-sensitivity interval is #258–#557. Published Freedom evidence spans 2 eval lineages and 1 of 1 components. Its strongest relative result is SALAD-Bench (attack_enhanced_information_safety_harms, #12 of 33); its weakest is SALAD-Bench (base_representation_toxicity, #32 of 33).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
ChineseSafescore#15 / 22↑53.67Source ↗official
CMoralEvalfamilial_morality#3 / 26↑0.56Source ↗official
CMoralEvalinternet_ethics#3 / 26↑0.54Source ↗official
CMoralEvalpersonal_morality#3 / 26↑0.52Source ↗official
CMoralEvalprofessional_ethics#3 / 26↑0.54Source ↗official
CMoralEvalsocial_morality#3 / 26↑0.54Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#97 / 270↑20.41Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#35 / 270↑93.33Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#144 / 270↑71.67Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#133 / 268↑96.41Source ↗official
FinEval Financial Security Knowledgefinancial_security_accuracy_pct#10 / 19↑73.1Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#15 / 33↑15.52Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#21 / 33↑7.17Source ↗official
SALAD-Benchattack_enhanced_malicious_use#15 / 33↑12.56Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#19 / 33↑9.38Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#20 / 33↑10.12Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#17 / 33↑12.55Source ↗official
SALAD-Benchbase_human_autonomy_integrity#3 / 33↑98.54Source ↗official
SALAD-Benchbase_information_safety_harms#11 / 33↑96.68Source ↗official
SALAD-Benchbase_malicious_use#2 / 33↑98.88Source ↗official
SALAD-Benchbase_misinformation_harms#2 / 33↑98.67Source ↗official
SALAD-Benchbase_representation_toxicity#2 / 33↑97.53Source ↗official
SALAD-Benchbase_socioeconomic_harms#3 / 33↑95.77Source ↗official
SALAD-Benchmcq_human_autonomy_integrity#8 / 33↑61.11Source ↗official
SALAD-Benchmcq_information_safety_harms#7 / 33↑51.11Source ↗official
SALAD-Benchmcq_malicious_use#6 / 33↑57.24Source ↗official
SALAD-Benchmcq_misinformation_harms#6 / 33↑59.05Source ↗official
SALAD-Benchmcq_representation_toxicity#6 / 33↑57.92Source ↗official
SALAD-Benchmcq_socioeconomic_harms#10 / 33↑46.11Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#234 / 270↓93.33Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#125 / 270↓71.67Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#136 / 268↓96.41Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#19 / 33↓15.52Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#12 / 33↓7.17Source ↗official
SALAD-Benchattack_enhanced_malicious_use#19 / 33↓12.56Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#14 / 33↓9.38Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#14 / 33↓10.12Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#17 / 33↓12.55Source ↗official
SALAD-Benchbase_human_autonomy_integrity#31 / 33↓98.54Source ↗official
SALAD-Benchbase_information_safety_harms#23 / 33↓96.68Source ↗official
SALAD-Benchbase_malicious_use#32 / 33↓98.88Source ↗official
SALAD-Benchbase_misinformation_harms#31 / 33↓98.67Source ↗official
SALAD-Benchbase_representation_toxicity#32 / 33↓97.53Source ↗official
SALAD-Benchbase_socioeconomic_harms#31 / 33↓95.77Source ↗official