← Evals

Evaluation profile

M3-SafetyBench

1sub-evals
0.139%total index weight
1components

Within-component eval weight: Misuse resistance 1.39%.

Model score (higher is better)Predicted score

About this eval

Multidimensional Chinese safety performance including red-team and refusal ability dimensions.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
overall_scorem3-safetybench/m3-safetybench.csv:overall_scoreMeasures safety knowledge and safe response selection across multilingual, multimodal, and multi-level risk scenarios.ordinary_harm_misuse_resistance:1.000m3-safetybench/m3-safetybench.csvHigher is better0.139%Misuse resistance 1.39%

overall_score

Measures safety knowledge and safe response selection across multilingual, multimodal, and multi-level risk scenarios.

RankModelValueRelative performanceProvenance
1doubao-pro-32k96.76official
2qwen-plus96.53official
3qwen-2.5-14b-instruct95.94official
4moonshot-v195.03official
5ernie-3.5-128k94.9official
6hunyuan-standard94.87official
7internlm2.5-20b-chat92.88official
8qwen-2.5-7b-instruct92.37official
9qwen2-7b-instruct91.7official
10internlm2.5-7b-chat89.33official
11deepseek-r1-distill-qwen-7b88.09official
12chatglm3-6b-32k87.27official
13baichuan-2-13b-chat86.39official
14qwen-2.5-1.5b-instruct85.87official
15glm-4-9b-chat84.2official
16baichuan-2-7b-chat82.91official
17qwen2-1.5b-instruct80.83official
18internlm2.5-1.8b-chat79.16official
19qwen-2.5-0.5b-instruct69.63official