← Evals

Evaluation profile

HELM Safety

5sub-evals
4.12%total index weight
3components

Within-component eval weight: Misuse resistance 13.5% · Fairness 21.3% · Benign helpfulness 6.39%.

Model score (higher is better)Predicted score

About this eval

Standardized HELM safety metrics spanning harmful compliance, bias, red-teaming, and over-refusal.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
anthropic_red_teamhelm/helm.csv:anthropic_red_teamMeasures whether the model resists adversarial conversations intended to provoke harmful, dishonest, or abusive responses.ordinary_harm_misuse_resistance:1.000helm-anthropic-red-teamHigher is better0.572%Misuse resistance 5.72%
bbqhelm/helm.csv:bbqMeasures whether the model avoids choosing stereotyped answers when questions about protected groups are ambiguous.fairness_nondiscrimination:1.000helm-bbqHigher is better2.13%Fairness 21.3%
harmbenchhelm/helm.csv:harmbenchMeasures whether the model avoids complying with requests designed to elicit harmful behavior.ordinary_harm_misuse_resistance:1.000harmbenchHigher is better0.143%Misuse resistance 1.43%
simple_safety_testshelm/helm.csv:simple_safety_testsMeasures whether the model gives safe responses to straightforward prompts about dangerous or harmful activities.ordinary_harm_misuse_resistance:1.000helm-simple-safety-testsHigher is better0.572%Misuse resistance 5.72%
xstesthelm/helm.csv:xstestMeasures whether the model answers harmless requests instead of refusing them merely because they contain superficially sensitive words.benign_helpfulness_non_overrefusal:0.556|ordinary_harm_misuse_resistance:0.444xstestHigher is better0.703%Benign helpfulness 6.39% · Misuse resistance 0.635%

anthropic_red_team

Measures whether the model resists adversarial conversations intended to provoke harmful, dishonest, or abusive responses.

RankModelValueRelative performanceProvenance
1claude-3-haiku1official
2gemini-1.5-flash0.999official
2gemini-1.5-pro0.999official
2granite-4.0-small-guardian0.999official
2qwen3-235b-a22b0.999official
2qwen3-next-80b-a3b0.999official
7claude-3-opus0.998official
7claude-3-sonnet0.998official
7claude-3.5-sonnet0.998official
10claude-3.7-sonnet0.997official
10gpt-4-turbo0.997official
10gpt-oss-20b0.997official
13gpt-4.1-nano0.996official
13gpt-5-nano0.996official
13granite-4.0-micro-guardian0.996official
13granite-4.0-small0.996official
13qwen2.5-72b-instruct0.996official
18gemini-2.5-pro0.995official
18gpt-5.10.995official
18gpt-oss-120b0.995official
18grok-3-mini0.995official
18palmyra-fin0.995official
23deepseek-llm-67b-chat0.994official
23gemini-2-flash0.994official
23gemini-2-flash-lite-preview0.994official
23gpt-4.5-preview0.994official
27gpt-4.10.993official
27gpt-4.1-mini0.993official
27kimi-k20.993official
27olmo-2-0325-32b-instruct0.993official
27palmyra-x50.993official
32gpt-4o0.991official
32gpt-50.991official
32gpt-5-mini0.991official
32granite-3.3-8b-instruct0.991official
32qwen2-72b-instruct0.991official
37glm-4.5-air0.99official
37qwen1.5-72b-chat0.99official
37qwen3-235b-a22b-fp8-throughput0.99official
40claude-opus-40.989official
41claude-sonnet-40.988official
41gemini-2.5-flash0.988official
41llama-3-8b-instruct0.988official
44gemini-2-pro-preview0.987official
44gemini-2.5-flash-lite0.987official
44granite-4.0-micro0.987official
47o3-mini0.986official
47palmyra-x-0040.986official
49qwen-2.5-7b-instruct0.985official
50gpt-4o-mini0.983official
50o10.983official
50o1-mini0.983official
50o30.983official
54claude-sonnet-4.50.982official
54llama-4-maverick0.982official
54o4-mini0.982official
57gpt-3.5-turbo0.981official
58command-r-plus0.98official
59palmyra-med0.978official
60deepseek-r10.975official
61deepseek-v30.971official
61gemini-3-pro-preview0.971official
63marin-8b-instruct0.97official
64claude-haiku-4.50.969official
65llama-3.1-8b-instruct0.968official
66llama-3-70b-instruct0.967official
67llama-3.1-405b-instruct0.965official
67llama-4-scout0.965official
69mistral-small-30.964official
70mixtral-8x22b-instruct0.962official
71grok-40.957official
72grok-30.955official
73command-r0.937official
74llama-3.1-70b-instruct0.932official
75mixtral-8x7b-instruct0.928official
76olmo-2-1124-13b-instruct0.907official
77olmoe-1b-7b-0125-instruct0.889official
78olmo-2-1124-7b-instruct0.88official
79mistral-7b-instruct0.806official
80dbrx-instruct0.766official

bbq

Measures whether the model avoids choosing stereotyped answers when questions about protected groups are ambiguous.

RankModelValueRelative performanceProvenance
1claude-sonnet-4.50.989official
2gpt-oss-120b0.985official
3gemini-3-pro-preview0.984official
4claude-opus-40.982official
5o30.979official
6glm-4.5-air0.978official
7gemini-2.5-flash0.977official
8gpt-5-nano0.976official
9o10.973official
10claude-sonnet-40.9725official
11o1-mini0.969official
12gpt-50.968official
13deepseek-v30.967official
13gpt-oss-20b0.967official
13grok-3-mini0.967official
16qwen3-235b-a22b-fp8-throughput0.966official
17deepseek-r10.9657official
18gemini-2.5-pro0.964official
19gpt-5-mini0.963official
20qwen3-235b-a22b0.962official
21gemini-2-pro-preview0.956official
22palmyra-x-0040.955official
23gemini-2-flash0.954official
23llama-3.1-70b-instruct0.954official
23qwen2.5-72b-instruct0.954official
26gpt-4o0.951official
26qwen2-72b-instruct0.951official
28claude-3.5-sonnet0.949official
28gemini-2.5-flash-lite0.949official
28kimi-k20.949official
31palmyra-x50.948official
32gemini-1.5-flash0.947official
33gemini-1.5-pro0.945official
33llama-3.1-405b-instruct0.945official
35palmyra-fin0.942official
36gpt-4-turbo0.941official
36o3-mini0.941official
38claude-3-opus0.94official
38o4-mini0.94official
40grok-40.937official
41grok-30.936official
42mistral-small-30.933official
43llama-4-maverick0.93official
44claude-haiku-4.50.928official
45gpt-4.10.926official
46granite-4.0-small0.925official
47claude-3.7-sonnet0.921official
47gpt-4.1-mini0.921official
49gemini-2-flash-lite-preview0.92official
49gpt-4.5-preview0.92official
51qwen3-next-80b-a3b0.917official
52mixtral-8x22b-instruct0.916official
53llama-3-70b-instruct0.91official
54qwen-2.5-7b-instruct0.906official
55claude-3-sonnet0.9official
55granite-4.0-small-guardian0.9official
57command-r-plus0.899official
58gpt-5.10.887official
59gpt-4o-mini0.882official
60gpt-4.1-nano0.875official
60llama-4-scout0.875official
62deepseek-llm-67b-chat0.862official
63mixtral-8x7b-instruct0.857official
64qwen1.5-72b-chat0.846official
65granite-3.3-8b-instruct0.84official
66palmyra-med0.801official
67dbrx-instruct0.792official
68llama-3.1-8b-instruct0.785official
69granite-4.0-micro0.779official
70llama-3-8b-instruct0.765official
71marin-8b-instruct0.755official
72granite-4.0-micro-guardian0.752official
73command-r0.724official
74olmo-2-0325-32b-instruct0.714official
75olmo-2-1124-13b-instruct0.704official
76gpt-3.5-turbo0.6513official
77claude-3-haiku0.625official
78mistral-7b-instruct0.6165official
79olmo-2-1124-7b-instruct0.525official
80olmoe-1b-7b-0125-instruct0.419official

harmbench

Measures whether the model avoids complying with requests designed to elicit harmful behavior.

RankModelValueRelative performanceProvenance
1gpt-oss-120b1official
2gpt-oss-20b0.987official
3gpt-5-nano0.984official
3o30.984official
5claude-3.5-sonnet0.981official
6gpt-50.976official
6gpt-5.10.976official
8claude-3-opus0.974official
8kimi-k20.974official
10gpt-5-mini0.971official
11o4-mini0.97official
12o10.963official
13claude-sonnet-40.9605official
14claude-haiku-4.50.959official
15claude-3-sonnet0.958official
15gpt-4.5-preview0.958official
17o3-mini0.952official
18claude-sonnet-4.50.92official
19gpt-4.10.917official
20claude-3-haiku0.913official
21claude-opus-40.904official
22gpt-4-turbo0.898official
23o1-mini0.885official
24gpt-4.1-nano0.868official
25granite-4.0-small-guardian0.864official
26granite-4.0-micro-guardian0.863official
27gpt-4.1-mini0.856official
28gpt-4o-mini0.849official
29qwen3-next-80b-a3b0.846official
30claude-3.7-sonnet0.843official
31olmo-2-0325-32b-instruct0.841official
31palmyra-fin0.841official
33gpt-4o0.829official
34granite-3.3-8b-instruct0.812official
35gemini-1.5-flash0.8official
36gemini-1.5-pro0.799official
37qwen3-235b-a22b0.796official
38qwen2-72b-instruct0.768official
39marin-8b-instruct0.751official
40palmyra-x-0040.738official
40palmyra-x50.738official
42qwen2.5-72b-instruct0.728official
43llama-3-8b-instruct0.727official
44gemini-3-pro-preview0.725official
45gemini-2-flash-lite-preview0.722official
46granite-4.0-small0.711official
47granite-4.0-micro0.695official
48qwen-2.5-7b-instruct0.677official
49gemini-2.5-flash-lite0.67official
50gpt-3.5-turbo0.6663official
51gemini-2-flash0.662official
52llama-4-maverick0.661official
53gemini-2.5-pro0.654official
54gemini-2-pro-preview0.653official
55deepseek-llm-67b-chat0.649official
56qwen1.5-72b-chat0.648official
57llama-3-70b-instruct0.64official
58olmoe-1b-7b-0125-instruct0.629official
59llama-3.1-405b-instruct0.627official
60gemini-2.5-flash0.626official
61llama-3.1-8b-instruct0.616official
62llama-4-scout0.6official
63olmo-2-1124-7b-instruct0.578official
64grok-3-mini0.572official
65qwen3-235b-a22b-fp8-throughput0.569official
66glm-4.5-air0.561official
67palmyra-med0.556official
68mixtral-8x22b-instruct0.521official
69olmo-2-1124-13b-instruct0.519official
70command-r0.504official
71deepseek-r10.4987official
72deepseek-v30.497official
73command-r-plus0.485official
74llama-3.1-70b-instruct0.469official
75grok-30.453official
76mistral-small-30.452official
77mixtral-8x7b-instruct0.451official
78grok-40.397official
79mistral-7b-instruct0.2885official
80dbrx-instruct0.271official

simple_safety_tests

Measures whether the model gives safe responses to straightforward prompts about dangerous or harmful activities.

RankModelValueRelative performanceProvenance
1claude-3-haiku1official
1claude-3-opus1official
1claude-3-sonnet1official
1claude-3.5-sonnet1official
1claude-3.7-sonnet1official
1claude-sonnet-4.51official
1command-r-plus1official
1gpt-4.11official
1gpt-4.1-mini1official
1gpt-4.5-preview1official
1gpt-5-mini1official
1gpt-5-nano1official
1gpt-oss-120b1official
1gpt-oss-20b1official
1granite-4.0-micro-guardian1official
1granite-4.0-small-guardian1official
1kimi-k21official
1o4-mini1official
1palmyra-fin1official
1palmyra-x-0041official
1palmyra-x51official
1qwen2.5-72b-instruct1official
1qwen3-235b-a22b1official
24gpt-50.998official
24gpt-5.10.998official
26claude-opus-40.9975official
26claude-sonnet-40.9975official
28qwen3-next-80b-a3b0.995official
29grok-3-mini0.993official
29llama-3-8b-instruct0.993official
29llama-4-maverick0.993official
32glm-4.5-air0.99official
32gpt-4-turbo0.99official
32gpt-4.1-nano0.99official
32llama-3-70b-instruct0.99official
32o10.99official
32o30.99official
32o3-mini0.99official
32qwen1.5-72b-chat0.99official
40claude-haiku-4.50.988official
40llama-3.1-405b-instruct0.988official
40llama-3.1-8b-instruct0.988official
43gemini-2-flash0.985official
43gpt-4o0.985official
43palmyra-med0.985official
43qwen2-72b-instruct0.985official
43qwen3-235b-a22b-fp8-throughput0.985official
48gemini-2.5-flash0.98official
48granite-4.0-small0.98official
48olmo-2-0325-32b-instruct0.98official
51deepseek-r10.9793official
52gpt-4o-mini0.978official
53gemini-2-flash-lite-preview0.977official
54gemini-1.5-pro0.975official
54gemini-2-pro-preview0.975official
54gemini-3-pro-preview0.975official
54granite-3.3-8b-instruct0.975official
58gemini-1.5-flash0.97official
58gemini-2.5-pro0.97official
58llama-4-scout0.97official
58o1-mini0.97official
62deepseek-llm-67b-chat0.968official
62grok-30.968official
64gemini-2.5-flash-lite0.965official
65qwen-2.5-7b-instruct0.96official
66marin-8b-instruct0.958official
67deepseek-v30.953official
68granite-4.0-micro0.945official
69command-r0.943official
70gpt-3.5-turbo0.9337official
71mistral-small-30.933official
72llama-3.1-70b-instruct0.925official
73grok-40.922official
74mixtral-8x22b-instruct0.915official
75mixtral-8x7b-instruct0.905official
76olmo-2-1124-13b-instruct0.81official
77olmo-2-1124-7b-instruct0.775official
78olmoe-1b-7b-0125-instruct0.725official
79mistral-7b-instruct0.6215official
80dbrx-instruct0.535official

xstest

Measures whether the model answers harmless requests instead of refusing them merely because they contain superficially sensitive words.

RankModelValueRelative performanceProvenance
1gemini-2.5-flash0.988official
1qwen3-235b-a22b-fp8-throughput0.988official
3gemini-2.5-pro0.987official
4glm-4.5-air0.986official
4qwen3-235b-a22b0.986official
6grok-3-mini0.984official
6palmyra-x-0040.984official
8qwen3-next-80b-a3b0.983official
9gpt-5.10.982official
9kimi-k20.982official
11gpt-4.10.979official
11qwen2.5-72b-instruct0.979official
13gemini-2.5-flash-lite0.978official
14gpt-4-turbo0.977official
14gpt-5-mini0.977official
16gpt-5-nano0.976official
17gpt-4.1-mini0.974official
17o4-mini0.974official
19gemini-3-pro-preview0.973official
19gpt-4o0.973official
19o30.973official
22deepseek-v30.971official
22gpt-50.971official
24claude-opus-40.97official
24o10.97official
24o1-mini0.97official
27claude-sonnet-40.969official
27qwen2-72b-instruct0.969official
29llama-3-70b-instruct0.968official
30grok-40.966official
30qwen-2.5-7b-instruct0.966official
32llama-4-maverick0.965official
33claude-3.7-sonnet0.964official
33claude-sonnet-4.50.964official
33grok-30.964official
36palmyra-med0.963official
37palmyra-fin0.962official
38deepseek-r10.9617official
39gpt-4.1-nano0.96official
39gpt-4o-mini0.96official
41llama-3.1-405b-instruct0.959official
42qwen1.5-72b-chat0.957official
43claude-3.5-sonnet0.956official
43llama-3-8b-instruct0.956official
43llama-4-scout0.956official
46gemini-2-pro-preview0.954official
46olmo-2-0325-32b-instruct0.954official
48gemini-2-flash0.953official
48llama-3.1-8b-instruct0.953official
48mistral-small-30.953official
48palmyra-x50.953official
52gpt-4.5-preview0.951official
53granite-4.0-small0.946official
54llama-3.1-70b-instruct0.945official
55mixtral-8x22b-instruct0.941official
55o3-mini0.941official
57command-r0.939official
58command-r-plus0.938official
59gpt-3.5-turbo0.9373official
60claude-haiku-4.50.932official
61mixtral-8x7b-instruct0.931official
62gemini-2-flash-lite-preview0.927official
62gpt-oss-120b0.927official
64olmo-2-1124-13b-instruct0.926official
65claude-3-opus0.925official
66gemini-1.5-flash0.921official
67marin-8b-instruct0.917official
68granite-3.3-8b-instruct0.907official
69gemini-1.5-pro0.904official
70deepseek-llm-67b-chat0.889official
71olmo-2-1124-7b-instruct0.888official
72gpt-oss-20b0.872official
73granite-4.0-micro0.869official
74claude-3-sonnet0.858official
75claude-3-haiku0.853official
76olmoe-1b-7b-0125-instruct0.843official
77mistral-7b-instruct0.8065official
78granite-4.0-small-guardian0.806official
79dbrx-instruct0.774official
80granite-4.0-micro-guardian0.754official