← Evals

Evaluation profile

BlueBench AttaQ-100

1sub-evals
0.136%total index weight
1components

Within-component eval weight: Misuse resistance 1.36%.

Model score (higher is better)Predicted score

About this eval

Learned harmlessness reward on 100 AttaQ harmful requests.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
attaq_harmlessness_reward_pctbluebench-attaq-100/bluebench-attaq-100.csv:attaq_harmlessness_reward_pctMeasures how strongly a learned reward model favors the assistant's responses to 100 harmful AttaQ requests.ordinary_harm_misuse_resistance:1.000bluebench-attaq-100Higher is better0.136%Misuse resistance 1.36%

attaq_harmlessness_reward_pct

Measures how strongly a learned reward model favors the assistant's responses to 100 harmful AttaQ requests.

RankModelValueRelative performanceProvenance
1granite-3.3-8b-instruct88.44official
2mistral-large88.04official
3granite-3.0-8b-instruct88.03official
4gpt-4o87.84official
5mistral-small-3.187.62official
6gpt-4.187.49official
7gpt-4.1-nano86.67official
8gpt-4.1-mini85.96official
9granite-3-2-2b-instruct84.87official
10mistral-medium84.68official
11llama-3.3-70b-instruct83.8official
12o4-mini83.55official
13llama-3.2-3b-instruct83.37official
14llama-3.2-1b-instruct83official
15o182.96official
16pixtral-12b82.45official
17llama-3.1-405b-instruct81.35official
18o3-mini80.61official