← Evals

Evaluation profile

RefusalBench

1sub-evals
0.945%total index weight
2components

Within-component eval weight: Benign helpfulness 8.4% · Misuse resistance 1.05%.

Model score (higher is better)Predicted score

About this eval

Refusal calibration on biological research prompts.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
youden_jrefusalbench/refusalbench.csv:youden_jMeasures whether the model refuses harmful prompts without unnecessarily refusing benign prompts.benign_helpfulness_non_overrefusal:0.500|ordinary_harm_misuse_resistance:0.500refusalbench/refusalbench.csvHigher is better0.945%Benign helpfulness 8.4% · Misuse resistance 1.05%

youden_j

Measures whether the model refuses harmful prompts without unnecessarily refusing benign prompts.

RankModelValueRelative performanceProvenance
1grok-4.200.9701official
2gemini-3.1-pro-preview0.9217official
3nemotron-3-super-120b-a12b0.8891official
4glm-50.7291official
5claude-sonnet-4.60.6766official
6claude-opus-4.50.6709official
7claude-opus-4.60.6667official
8minimax-m2.70.6604official
9qwen3-next-80b-a3b0.4545official
10gemini-3.1-flash-lite0.4533official
11gpt-5.50.4213official
12claude-opus-4.70.234official
13deepseek-r10.1424official
14nova-pro0.09333official
15kimi-k2.60.08511official
16gpt-5.4-mini0.06411official
17deepseek-v3.20.01333official
17llama-3.3-70b-instruct0.01333official
19mistral-large-30.00906official