Evaluation profile
JailBench
1sub-evals
0.12%total index weight
1components
Within-component eval weight: Misuse resistance 1.2%.
Model score (lower is better)Predicted score
About this eval
Jailbreak susceptibility across Chinese safety categories.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| jailbreak_success_ratejailbench/jailbench.csv:jailbreak_success_rateMeasures whether jailbreak attacks can make the model comply with harmful requests. | ordinary_harm_misuse_resistance:1.000jailbench/jailbench.csv | Lower is better | 0.12% | Misuse resistance 1.2% |
jailbreak_success_rate
Measures whether jailbreak attacks can make the model comply with harmful requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-4 | 34.5 | official | |
| 2 | llama-3-8b-instruct | 43.94 | official | |
| 3 | llama-2-7b-chat | 48.68 | official | |
| 4 | qwen2-7b-chat | 49.73 | official | |
| 5 | internlm2-chat-7b | 51.22 | official | |
| 6 | llama-2-13b-chat | 55.39 | official | |
| 7 | chatglm3 | 58.81 | official | |
| 8 | vicuna-7b-v1.5 | 59.27 | official | |
| 9 | vicuna-13b-v1.5 | 66.32 | official | |
| 10 | deepseek-v3 | 70.67 | official | |
| 11 | qwen1.5-7b-chat | 71.6 | official | |
| 12 | gpt-3.5-turbo | 73.86 | official | |
| 13 | glm-4-9b-chat | 74.95 | official | |
| 14 | mistral-7b-instruct | 78.14 | official |