← Models

Model profile

Internlm Chat 7B

InternLMdeveloper
2023-07-06release date
#148 / 267overall rank
5eval lineages

Evidence summary

Internlm Chat 7B has an estimated overall rank of #148; its 90% source-sensitivity interval is #39–#215. Its behavior-only rank is #147; company governance moves the combined estimate to #148. Published evidence spans 5 evals and 4 of 7 behavior components. Its strongest relative result is FLAMES (legality, #1 of 13); its weakest is SALAD-Bench (mcq_representation_toxicity, #30 of 33).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
Fake Alignment (FINE)multiple_choice_safe_decision_rate#6 / 1457.33↑ higherSource ↗official
Fake Alignment (FINE)open_ended_safe_response_rate#11 / 1492↑ higherSource ↗official
FLAMESdata_protection#2 / 1361.84↑ higherSource ↗official
FLAMESfairness#2 / 1344.58↑ higherSource ↗official
FLAMESlegality#1 / 1376.09↑ higherSource ↗official
FLAMESmorality#3 / 1351.24↑ higherSource ↗official
FLAMESsafety#5 / 1335.9↑ higherSource ↗official
SafetyBenchEM#6 / 2175.4↑ higherSource ↗official
SafetyBenchIA#8 / 2179.5↑ higherSource ↗official
SafetyBenchMH#6 / 2184.3↑ higherSource ↗official
SafetyBenchOFF#12 / 2167.2↑ higherSource ↗official
SafetyBenchPH#7 / 2174.15↑ higherSource ↗official
SafetyBenchPP#7 / 2178.7↑ higherSource ↗official
SafetyBenchUB#8 / 2164.75↑ higherSource ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#11 / 3321.12↑ higherSource ↗official
SALAD-Benchattack_enhanced_information_safety_harms#16 / 3316.61↑ higherSource ↗official
SALAD-Benchattack_enhanced_malicious_use#12 / 3322.51↑ higherSource ↗official
SALAD-Benchattack_enhanced_misinformation_harms#11 / 3319.9↑ higherSource ↗official
SALAD-Benchattack_enhanced_representation_toxicity#9 / 3323.71↑ higherSource ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#11 / 3322.08↑ higherSource ↗official
SALAD-Benchbase_human_autonomy_integrity#10 / 3396.85↑ higherSource ↗official
SALAD-Benchbase_information_safety_harms#15 / 3395.13↑ higherSource ↗official
SALAD-Benchbase_malicious_use#12 / 3396.28↑ higherSource ↗official
SALAD-Benchbase_misinformation_harms#5 / 3397.05↑ higherSource ↗official
SALAD-Benchbase_representation_toxicity#7 / 3394.37↑ higherSource ↗official
SALAD-Benchbase_socioeconomic_harms#15 / 3390.95↑ higherSource ↗official
SALAD-Benchmcq_human_autonomy_integrity#30 / 330↑ higherSource ↗official
SALAD-Benchmcq_information_safety_harms#30 / 330↑ higherSource ↗official
SALAD-Benchmcq_malicious_use#30 / 330.0641↑ higherSource ↗official
SALAD-Benchmcq_misinformation_harms#30 / 330↑ higherSource ↗official
SALAD-Benchmcq_representation_toxicity#30 / 330.1042↑ higherSource ↗official
SALAD-Benchmcq_socioeconomic_harms#30 / 330↑ higherSource ↗official
SuperCLUE Safetyinstruction_attack#20 / 3158.62↑ higherSource ↗official
SuperCLUE Safetyresponsible_ai#28 / 3145.45↑ higherSource ↗official
SuperCLUE Safetytraditional_safety#28 / 3165.96↑ higherSource ↗official