← Evals

Evaluation profile

SORRY-Bench

1sub-evals
0.245%total index weight
1components

Within-component eval weight: Misuse resistance 2.45%.

Model score (lower is better)Predicted score

About this eval

Refusal or harmful-compliance behavior across unsafe topics.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
avgsorrybench/sorrybench.csv:avgMeasures whether the model refuses unsafe requests across many policy categories rather than providing harmful assistance.ordinary_harm_misuse_resistance:1.000sorrybenchLower is better0.245%Misuse resistance 2.45%

avg

Measures whether the model refuses unsafe requests across many policy categories rather than providing harmful assistance.

RankModelValueRelative performanceProvenance
1claude-2.10.05official
2claude-20.07official
2claude-3-opus0.07official
2gemini-1.5-pro0.07official
5claude-3-haiku0.08official
5claude-instant-1.20.08official
5gemini-1.5-flash0.08official
8gemma-2-9b-it0.1official
9claude-3-sonnet0.11official
9gemma-2-27b-it0.11official
11llama-2-70b-chat0.12official
12claude-3.5-sonnet0.14official
12llama-2-7b-chat0.14official
12llama-3.1-8b-instruct0.14official
15llama-2-13b-chat0.15official
16gemma-7b-it0.18official
17gpt-3.5-turbo0.18official
18gemma-2b-it0.19official
19llama-3-8b-instruct0.22official
20llama-3.1-405b-instruct0.24official
20o1-mini0.24official
22o1-preview0.25official
23gpt-4-turbo0.2533official
24qwen1.5-4b-chat0.27official
25qwen1.5-32b-chat0.28official
26gpt-40.29official
27gpt-4o0.3official
28vicuna-13b-v1.50.32official
29gemini-1.0-pro0.33official
30qwen1.5-14b-chat0.34official
31aurora-m0.35official
31llama-3-70b-instruct0.35official
31vicuna-7b-v1.50.35official
34chatglm3-6b0.36official
34gpt-4o-mini0.36official
34qwen1.5-72b-chat0.36official
34zephyr-7b-r2d20.36official
38llama-3.1-70b-instruct0.39official
39qwen1.5-7b-chat0.4official
40yi-1.5-34b-chat0.41official
41qwen1.5-0.5b-chat0.43official
42yi-1.5-6b-chat0.49official
43mixtral-8x7b-instruct0.56official
43vicuna-33b-v1.30.56official
45mistral-large-20.6official
46openchat-3.50.68official
47qwen1.5-1.8b-chat0.74official
48mistral-7b-instruct0.785official
49dolphin-2.6-mixtral-8x7b0.85official
50zephyr-7b-beta0.87official
51dolphin-2.2.1-mistral-7b0.91official