← Evals

Evaluation profile

HarmBench

1sub-evals
0.0846%total index weight
1components

Within-component eval weight: Misuse resistance 0.846%.

Model score (lower is better)Predicted score

About this eval

Harmful compliance or attack success under harmful request benchmarks.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
drharmbench/harmbench.csv:drMeasures whether the model refuses direct requests for assistance with harmful behavior.ordinary_harm_misuse_resistance:1.000harmbenchLower is better0.0846%Misuse resistance 0.846%

dr

Measures whether the model refuses direct requests for assistance with harmful behavior.

RankModelValueRelative performanceProvenance
1llama-2-7b-chat0.8official
2claude-22official
2claude-2.12official
4llama-2-13b-chat2.8official
4llama-2-70b-chat2.8official
6claude-15official
7gpt-4-turbo9.3official
8qwen-7b-chat13official
9r2d214.2official
10qwen-14b-chat16.5official
11gemini-1.0-pro18official
12qwen-72b-chat18.3official
13baichuan-2-7b-chat18.8official
14baichuan-2-13b-chat19.3official
15vicuna-13b-v1.519.8official
16gpt-421official
17vicuna-7b-v1.524.3official
18gpt-3.5-turbo27.15official
19koala-13b27.3official
20koala-7b38.3official
21orca-2-7b39official
22orca-2-13b44.5official
23openchat-3.546official
24mistral-7b-instruct46.3official
25mixtral-8x7b-instruct47.3official
26starling-7b57official
27solar-10.7b-instruct61.3official
28zephyr-7b-beta65.8official