← Models

Model profile

Chatglm2 6B

Z.aideveloper
2023-06-25release date
#225 / 267overall rank
6eval lineages

Evidence summary

Chatglm2 6B has an estimated overall rank of #225; its 90% source-sensitivity interval is #157–#250. Its behavior-only rank is #230; company governance moves the combined estimate to #225. Published evidence spans 6 evals and 3 of 7 behavior components. Its strongest relative result is SuperCLUE Safety (traditional_safety, #10 of 31); its weakest is Do-Not-Answer (human_harmlessness_rate, #6 of 6).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
ChiSafetyBenchharmful_response_rate#12 / 141.08↓ lowerSource ↗official
Do-Not-Answerhuman_harmlessness_rate#6 / 690.95↑ higherSource ↗official
Fake Alignment (FINE)multiple_choice_safe_decision_rate#13 / 1417.33↑ higherSource ↗official
Fake Alignment (FINE)open_ended_safe_response_rate#13 / 1485.33↑ higherSource ↗official
FLAMESdata_protection#6 / 1343.42↑ higherSource ↗official
FLAMESfairness#10 / 1331.73↑ higherSource ↗official
FLAMESlegality#12 / 1328.26↑ higherSource ↗official
FLAMESmorality#8 / 1343.28↑ higherSource ↗official
FLAMESsafety#11 / 1322.61↑ higherSource ↗official
SafetyBenchEM#10 / 2169.4↑ higherSource ↗official
SafetyBenchIA#10 / 2178.2↑ higherSource ↗official
SafetyBenchMH#10 / 2182↑ higherSource ↗official
SafetyBenchOFF#10 / 2168.1↑ higherSource ↗official
SafetyBenchPH#12 / 2167.9↑ higherSource ↗official
SafetyBenchPP#11 / 2176↑ higherSource ↗official
SafetyBenchUB#12 / 2161.6↑ higherSource ↗official
SuperCLUE Safetyinstruction_attack#20 / 3158.62↑ higherSource ↗official
SuperCLUE Safetyresponsible_ai#24 / 3147.27↑ higherSource ↗official
SuperCLUE Safetytraditional_safety#10 / 3178.72↑ higherSource ↗official