Evaluation profile
SafetyPrompts
1sub-evals
0.316%total index weight
4components
Within-component eval weight: Human rights 1.08% · Fairness 0.439% · Truthfulness 0.198% · Misuse resistance 0.796%.
Model score (higher is better)Predicted score
About this eval
Safety performance on Chinese safety prompts.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| scoresafety-prompts/safety-prompts.csv:scoreMeasures whether responses avoid offensive, biased, illegal, physically harmful, psychologically harmful, privacy-invasive, and unethical content under varied instruction attacks. | human_rights_systemic_harm:0.214|fairness_nondiscrimination:0.071|truthfulness_honesty:0.071|ordinary_harm_misuse_resistance:0.643safety-prompts/safety-prompts.csv | Higher is better | 0.316% | Human rights 1.08% · Fairness 0.439% · Truthfulness 0.198% · Misuse resistance 0.796% |
score
Measures whether responses avoid offensive, biased, illegal, physically harmful, psychologically harmful, privacy-invasive, and unethical content under varied instruction attacks.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-3.5-turbo | 98.37 | official | |
| 2 | chatglm-130b | 96.91 | official | |
| 3 | chatglm-6b | 96.81 | official | |
| 4 | minichat-100b | 95.25 | official | |
| 5 | minichat-10b | 92.18 | official | |
| 6 | belle | 86.14 | official | |
| 7 | text-davinci-003 | 84.76 | official | |
| 8 | text-davinci-002 | 59.89 | official | |
| 9 | text-davinci-001 | 57.92 | official | |
| 10 | chatyuan | 48.9 | official | |
| 11 | opd | 46.74 | official | |
| 12 | davinci | 42.98 | official | |
| 13 | bloom | 41.65 | official | |
| 14 | eva2.0 | 21.68 | official | |
| 15 | eva1.0 | 19.7 | official |