← Evals

Evaluation profile

OR-Bench

2sub-evals
1.45%total index weight
2components

Within-component eval weight: Benign helpfulness 12.9% · Misuse resistance 1.6%.

Model score (lower is better)Predicted score

About this eval

Appropriate refusal calibration across seemingly toxic but safe requests and genuinely toxic requests.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
over_refusal_rateor-bench/or-bench.csv:over_refusal_rateMeasures whether the model incorrectly refuses seemingly sensitive but safe requests.benign_helpfulness_non_overrefusal:1.000or-bench/or-bench.csvLower is better1.29%Benign helpfulness 12.9%
toxic_acceptance_rateor-bench/or-bench.csv:toxic_acceptance_rateMeasures whether the model answers genuinely harmful requests instead of refusing them.ordinary_harm_misuse_resistance:1.000or-bench/or-bench.csvLower is better0.16%Misuse resistance 1.6%

over_refusal_rate

Measures whether the model incorrectly refuses seemingly sensitive but safe requests.

RankModelValueRelative performanceProvenance
1gpt-4o6.8official
2gemini-1.0-pro9.7official
3mistral-large9.8official
4gpt-4-turbo12.5official
5mistral-small13.3official
6mistral-medium14official
7gemma-7b-it26.3official
8gpt-3.5-turbo36.17official
9llama-3-70b-instruct37.7official
10qwen1.5-7b-chat39.2official
11claude-3.5-sonnet43.8official
12qwen1.5-72b-chat46.9official
13qwen1.5-32b-chat50.8official
14gemma-2-27b-it62official
15llama-3-8b-instruct69.4official
16gemma-2-9b-it79.9official
17gemini-1.5-flash84.3official
18llama-2-7b-chat87.5official
19gemini-1.5-pro88official
20claude-3-opus91official
20llama-2-13b-chat91official
22claude-3-sonnet94.5official
23llama-2-70b-chat96.1official
24claude-3-haiku96.3official
25claude-2.199.8official

toxic_acceptance_rate

Measures whether the model answers genuinely harmful requests instead of refusing them.

RankModelValueRelative performanceProvenance
1claude-2.10official
2claude-3-haiku0.3official
2claude-3-sonnet0.3official
2llama-2-13b-chat0.3official
2llama-2-70b-chat0.3official
6llama-2-7b-chat0.4official
7gemini-1.5-pro0.6official
8gemini-1.5-flash1.2official
9gemma-2-9b-it1.5official
10claude-3-opus1.9official
11gemma-2-27b-it3.2official
12claude-3.5-sonnet3.4official
13qwen1.5-32b-chat4.4official
14llama-3-8b-instruct5official
15gpt-4-turbo5.25official
16qwen1.5-72b-chat5.6official
17gemma-7b-it14.5official
18qwen1.5-7b-chat15official
19gpt-4o15.1official
20gpt-3.5-turbo17.03official
21mistral-small20.3official
22gemini-1.0-pro21.3official
22llama-3-70b-instruct21.3official
24mistral-medium22.5official
25mistral-large27.2official