Evaluation profile
BlueBench AttaQ-100
1sub-evals
0.136%total index weight
1components
Within-component eval weight: Misuse resistance 1.36%.
Model score (higher is better)Predicted score
About this eval
Learned harmlessness reward on 100 AttaQ harmful requests.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| attaq_harmlessness_reward_pctbluebench-attaq-100/bluebench-attaq-100.csv:attaq_harmlessness_reward_pctMeasures how strongly a learned reward model favors the assistant's responses to 100 harmful AttaQ requests. | ordinary_harm_misuse_resistance:1.000bluebench-attaq-100 | Higher is better | 0.136% | Misuse resistance 1.36% |
attaq_harmlessness_reward_pct
Measures how strongly a learned reward model favors the assistant's responses to 100 harmful AttaQ requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | granite-3.3-8b-instruct | 88.44 | official | |
| 2 | mistral-large | 88.04 | official | |
| 3 | granite-3.0-8b-instruct | 88.03 | official | |
| 4 | gpt-4o | 87.84 | official | |
| 5 | mistral-small-3.1 | 87.62 | official | |
| 6 | gpt-4.1 | 87.49 | official | |
| 7 | gpt-4.1-nano | 86.67 | official | |
| 8 | gpt-4.1-mini | 85.96 | official | |
| 9 | granite-3-2-2b-instruct | 84.87 | official | |
| 10 | mistral-medium | 84.68 | official | |
| 11 | llama-3.3-70b-instruct | 83.8 | official | |
| 12 | o4-mini | 83.55 | official | |
| 13 | llama-3.2-3b-instruct | 83.37 | official | |
| 14 | llama-3.2-1b-instruct | 83 | official | |
| 15 | o1 | 82.96 | official | |
| 16 | pixtral-12b | 82.45 | official | |
| 17 | llama-3.1-405b-instruct | 81.35 | official | |
| 18 | o3-mini | 80.61 | official |