← Models

Model profile

Baichuan 13B Chat

2023-07-08release date
2eval lineages

Evidence summary

Published evidence spans 2 evals and 3 of 7 behavior components. Its strongest relative result is SuperCLUE Safety (traditional_safety, #5 of 31); its weakest is SuperCLUE Safety (responsible_ai, #31 of 31).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
SafetyBenchEM#11 / 2168.4↑ higherSource ↗official
SafetyBenchIA#9 / 2178.65↑ higherSource ↗official
SafetyBenchMH#8 / 2183.15↑ higherSource ↗official
SafetyBenchOFF#15 / 2159.25↑ higherSource ↗official
SafetyBenchPH#11 / 2168.2↑ higherSource ↗official
SafetyBenchPP#9 / 2177↑ higherSource ↗official
SafetyBenchUB#10 / 2162.65↑ higherSource ↗official
SuperCLUE Safetyinstruction_attack#29 / 3151.72↑ higherSource ↗official
SuperCLUE Safetyresponsible_ai#31 / 3140↑ higherSource ↗official
SuperCLUE Safetytraditional_safety#5 / 3180.85↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

CCP-aligned censorship behavior

DimensionValueDistribution
Political-question refusal rate (ZH/EN mean)49.5
Completely inaccurate rate8.32