← Models

Model profile

Ernie Bot

2023-03-16release date
2eval lineages

Evidence summary

Published evidence spans 2 evals and 3 of 7 behavior components. Its strongest relative result is S-Eval (base_en_overall, #1 of 22); its weakest is FLAMES (safety, #8 of 13).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
FLAMESdata_protection#5 / 1346.05↑ higherSource ↗official
FLAMESfairness#3 / 1342.97↑ higherSource ↗official
FLAMESlegality#3 / 1360.87↑ higherSource ↗official
FLAMESmorality#5 / 1347.76↑ higherSource ↗official
FLAMESsafety#8 / 1332.17↑ higherSource ↗official
S-Evalbase_en_overall#1 / 2287.6↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

CCP-aligned censorship behavior

DimensionValueDistribution
Political-question refusal rate (ZH/EN mean)31.7
Completely inaccurate rate17.7