← Evals

Evaluation profile

AIRBench 2024 Safety Scenarios

1sub-evals
0.429%total index weight
1components

Within-component eval weight: Misuse resistance 4.29%.

Model score (higher is better)Predicted score

About this eval

AIRBench 2024 safety-scenarios performance.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
safety_scenarioshelm/airbench2024.csv:safety_scenariosMeasures whether the model responds safely across diverse real-world risk scenarios generated and audited for AIRBench.ordinary_harm_misuse_resistance:1.000helm/airbench2024.csvHigher is better0.429%Misuse resistance 4.29%

safety_scenarios

Measures whether the model responds safely across diverse real-world risk scenarios generated and audited for AIRBench.

RankModelValueRelative performanceProvenance
1claude-haiku-4.50.932official
2claude-sonnet-4.50.898official
3claude-3.5-sonnet0.8835official
4claude-sonnet-40.883official
5gpt-oss-120b0.88official
6gpt-5-nano0.878official
7gpt-50.877official
8qwen3-next-80b-a3b0.867official
9gpt-5.10.862official
10gpt-oss-20b0.86official
11claude-opus-40.857official
11gpt-5-mini0.857official
13claude-3-sonnet0.847official
14o30.845official
15claude-3-opus0.844official
16claude-3-haiku0.827official
17granite-3.3-8b-instruct-guardian0.825official
18granite-4.0-small-guardian0.821official
19claude-3.7-sonnet0.818official
20granite-4.0-micro-guardian0.804official
21o10.8official
22qwen3-235b-a22b0.79official
23o4-mini0.785official
24palmyra-x50.782official
25granite-3.3-8b-instruct0.76official
26gemini-1.5-pro0.7505official
27o3-mini0.749official
28gpt-4.5-preview0.741official
28kimi-k20.741official
30gemini-2.5-pro0.736official
31gemini-1.5-flash0.7325official
32gemini-3-pro-preview0.732official
33gpt-4-turbo0.719official
34granite-4.0-small0.716official
35llama-3-8b-instruct0.709official
36gemini-2.5-flash0.687official
37llama-4-maverick0.686official
38gemini-2-pro-preview0.684official
39gemini-2-flash-lite-preview0.675official
40palmyra-fin0.663official
41gemini-2-flash0.662official
42granite-4.0-micro0.661official
43gemini-2.5-flash-lite0.658official
44gpt-4.10.648official
45llama-3-70b-instruct0.646official
46gpt-40.642official
47llama-3.1-8b-instruct0.623official
48qwen2-72b-instruct0.621official
49gpt-4.1-nano0.615official
50gpt-4.1-mini0.604official
51qwen2.5-72b-instruct0.59official
52llama-3.1-405b-instruct0.586official
53gemini-1.0-pro0.582official
54palmyra-med0.578official
55gpt-4o0.5755official
56glm-4.5-air0.571official
57gpt-4o-mini0.563official
58qwen3-235b-a22b-fp8-throughput0.56official
59gpt-3.5-turbo0.5577official
60yi-34b-chat0.536official
61grok-3-mini0.535official
62deepseek-r10.529official
63llama-4-scout0.523official
64grok-30.513official
65deepseek-llm-67b-chat0.505official
66qwen1.5-72b-chat0.486official
67qwen-2.5-7b-instruct0.47official
68o1-mini0.452official
69grok-40.444official
70palmyra-x-0040.442official
71mixtral-8x22b-instruct0.44official
72llama-3.1-70b-instruct0.425official
73deepseek-v30.408official
74mixtral-8x7b-instruct0.391official
75mistral-large-20.353official
76mistral-small-30.328official
77mistral-7b-instruct0.326official
78command-r0.318official
79command-r-plus0.293official
80dbrx-instruct0.254official