Evaluation profile
ChineseSafe
1sub-evals
0.15%total index weight
1components
Within-component eval weight: Misuse resistance 1.5%.
Model score (higher is better)Predicted score
About this eval
Accuracy at classifying safe versus unsafe Chinese content under the paper's perplexity-based evaluation.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| scorechinesesafe/chinesesafe.csv:scoreMeasures whether the model correctly classifies Chinese prompts and responses as safe or unsafe using the benchmark’s perplexity method. | ordinary_harm_misuse_resistance:1.000chinesesafe/chinesesafe.csv | Higher is better | 0.15% | Misuse resistance 1.5% |
score
Measures whether the model correctly classifies Chinese prompts and responses as safe or unsafe using the benchmark’s perplexity method.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | baichuan-2-13b-chat | 70.43 | official | |
| 2 | deepseek-llm-67b-chat | 68.08 | official | |
| 3 | yi-1.5-34b-chat | 66.02 | official | |
| 4 | gemma-1.1-7b-it | 64.32 | official | |
| 5 | qwen1.5-72b-chat | 63.67 | official | |
| 6 | qwen1.5-7b-chat | 62.48 | official | |
| 7 | qwen1.5-14b-chat | 61.29 | official | |
| 8 | yi-1.5-9b-chat | 60.35 | official | |
| 9 | opt-66b | 59.93 | official | |
| 10 | deepseek-llm-7b-chat | 56.79 | official | |
| 11 | gpt-j-6b | 55.98 | official | |
| 12 | ziya2-13b-chat | 55.25 | official | |
| 13 | baichuan-2-7b-chat | 53.99 | official | |
| 14 | opt-30b | 53.82 | official | |
| 15 | internlm2-chat-20b | 53.67 | official | |
| 16 | glm-4-9b-chat | 50.03 | official | |
| 17 | internlm2-chat-7b | 49.49 | official | |
| 18 | opt-13b | 49.31 | official | |
| 19 | opt-6.7b | 48.54 | official | |
| 20 | mistral-7b-instruct | 42.99 | official | |
| 21 | llama3-chatqa-1.5-8b | 42.11 | official | |
| 22 | llama3-chatqa-1.5-70b | 40.41 | official |