← Evals

Evaluation profile

SafetyPrompts

1sub-evals
0.316%total index weight
4components

Within-component eval weight: Human rights 1.08% · Fairness 0.439% · Truthfulness 0.198% · Misuse resistance 0.796%.

Model score (higher is better)Predicted score

About this eval

Safety performance on Chinese safety prompts.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
scoresafety-prompts/safety-prompts.csv:scoreMeasures whether responses avoid offensive, biased, illegal, physically harmful, psychologically harmful, privacy-invasive, and unethical content under varied instruction attacks.human_rights_systemic_harm:0.214|fairness_nondiscrimination:0.071|truthfulness_honesty:0.071|ordinary_harm_misuse_resistance:0.643safety-prompts/safety-prompts.csvHigher is better0.316%Human rights 1.08% · Fairness 0.439% · Truthfulness 0.198% · Misuse resistance 0.796%

score

Measures whether responses avoid offensive, biased, illegal, physically harmful, psychologically harmful, privacy-invasive, and unethical content under varied instruction attacks.

RankModelValueRelative performanceProvenance
1gpt-3.5-turbo98.37official
2chatglm-130b96.91official
3chatglm-6b96.81official
4minichat-100b95.25official
5minichat-10b92.18official
6belle86.14official
7text-davinci-00384.76official
8text-davinci-00259.89official
9text-davinci-00157.92official
10chatyuan48.9official
11opd46.74official
12davinci42.98official
13bloom41.65official
14eva2.021.68official
15eva1.019.7official