← Models

Model profile

Internlm2 Chat 7B

InternLMdeveloper
2024-01-17release date
#52 / 267overall rank
4eval lineages

Evidence summary

Internlm2 Chat 7B has an estimated overall rank of #52; its 90% source-sensitivity interval is #14–#221. Its behavior-only rank is #54; company governance moves the combined estimate to #52. Published evidence spans 4 evals and 4 of 7 behavior components. Its strongest relative result is SALAD-Bench (base_misinformation_harms, #2 of 33); its weakest is ChineseSafe (score, #17 of 22).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
ChineseSafescore#17 / 2249.49↑ higherSource ↗official
CMoralEvalfamilial_morality#5 / 260.51↑ higherSource ↗official
CMoralEvalinternet_ethics#4 / 260.51↑ higherSource ↗official
CMoralEvalpersonal_morality#4 / 260.5↑ higherSource ↗official
CMoralEvalprofessional_ethics#4 / 260.52↑ higherSource ↗official
CMoralEvalsocial_morality#4 / 260.51↑ higherSource ↗official
JailBenchjailbreak_success_rate#5 / 1451.22↓ lowerSource ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#11 / 3321.12↑ higherSource ↗official
SALAD-Benchattack_enhanced_information_safety_harms#14 / 3317.59↑ higherSource ↗official
SALAD-Benchattack_enhanced_malicious_use#8 / 3324.31↑ higherSource ↗official
SALAD-Benchattack_enhanced_misinformation_harms#16 / 3316.45↑ higherSource ↗official
SALAD-Benchattack_enhanced_representation_toxicity#13 / 3319.27↑ higherSource ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#12 / 3320.35↑ higherSource ↗official
SALAD-Benchbase_human_autonomy_integrity#7 / 3398.02↑ higherSource ↗official
SALAD-Benchbase_information_safety_harms#18 / 3394.45↑ higherSource ↗official
SALAD-Benchbase_malicious_use#3 / 3398.63↑ higherSource ↗official
SALAD-Benchbase_misinformation_harms#2 / 3398.67↑ higherSource ↗official
SALAD-Benchbase_representation_toxicity#3 / 3397.19↑ higherSource ↗official
SALAD-Benchbase_socioeconomic_harms#5 / 3394.71↑ higherSource ↗official
SALAD-Benchmcq_human_autonomy_integrity#5 / 3363.89↑ higherSource ↗official
SALAD-Benchmcq_information_safety_harms#5 / 3357.5↑ higherSource ↗official
SALAD-Benchmcq_malicious_use#5 / 3362.24↑ higherSource ↗official
SALAD-Benchmcq_misinformation_harms#5 / 3362.62↑ higherSource ↗official
SALAD-Benchmcq_representation_toxicity#5 / 3361.25↑ higherSource ↗official
SALAD-Benchmcq_socioeconomic_harms#5 / 3356.11↑ higherSource ↗official