← Models

Model profile

Chatglm 6B

Z.aideveloper
2023-03-14release date
#199 / 267overall rank
5eval lineages

Evidence summary

Chatglm 6B has an estimated overall rank of #199; its 90% source-sensitivity interval is #102–#234. Its behavior-only rank is #198; company governance moves the combined estimate to #199. Published evidence spans 5 evals and 4 of 7 behavior components. Its strongest relative result is SafetyPrompts (score, #3 of 15); its weakest is ChiSafetyBench (harmful_response_rate, #14 of 14).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
Chinese Bias Benchmark for Question Answeringbias_score#7 / 100.4437↓ lowerSource ↗official
ChiSafetyBenchharmful_response_rate#14 / 141.95↓ lowerSource ↗official
CValuesscore#3 / 1095↑ higherSource ↗official
FLAMESdata_protection#11 / 1332.89↑ higherSource ↗official
FLAMESfairness#12 / 1326.91↑ higherSource ↗official
FLAMESlegality#5 / 1350↑ higherSource ↗official
FLAMESmorality#9 / 1340.3↑ higherSource ↗official
FLAMESsafety#12 / 1315.38↑ higherSource ↗official
SafetyPromptsscore#3 / 1596.81↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

CCP-aligned censorship behavior

DimensionValueDistribution
Political-question refusal rate (ZH/EN mean)7.61
Completely inaccurate rate8.65