← Evals

Evaluation profile

DSPSafeBench

1sub-evals
0.222%total index weight
1components

Within-component eval weight: Misuse resistance 2.22%.

Model score (higher is better)Predicted score

About this eval

Aggregate compliance rate on adversarial Chinese content-safety prompts.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
scoredspsafebench/dspsafebench.csv:scoreMeasures whether the model produces responses that comply with the benchmark’s safety criteria across diverse Chinese prompts.ordinary_harm_misuse_resistance:1.000dspsafebench/dspsafebench.csvHigher is better0.222%Misuse resistance 2.22%

score

Measures whether the model produces responses that comply with the benchmark’s safety criteria across diverse Chinese prompts.

RankModelValueRelative performanceProvenance
1yi-1.5-9b-chat-16k79.37official
2phi-3-mini-4k-instruct78.62official
3internlm2.5-7b-chat77.64official
3minicpm3-4b77.64official
5qwen-2.5-7b-instruct73.51official
6mistral-7b-instruct73.04official
7glm-4-9b-chat72.43official
8gemma-2-9b-it72.34official
9qwen-2.5-1.5b-instruct71.26official
10gemma-2-2b-it70.42official
11baichuan-2-7b-chat65.31official
12llama-3.1-8b-instruct61.51official