← Evals

Evaluation profile

VETO Misfired Alignment

1sub-evals
1.93%total index weight
2components

Within-component eval weight: Fairness 10.3% · Benign helpfulness 9%.

Model score (lower is better)Predicted score

About this eval

VETO Misfired Alignment

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
misfired_alignment_rate_pctveto/veto.csv:misfired_alignment_rate_pctMeasures misfired alignment rate pct in the cited first-party or paper evaluation.fairness_nondiscrimination:0.650|benign_helpfulness_non_overrefusal:0.350veto-misfired-alignmentLower is better1.93%Fairness 10.3% · Benign helpfulness 9%

misfired_alignment_rate_pct

Measures misfired alignment rate pct in the cited first-party or paper evaluation.

RankModelValueRelative performanceProvenance
1deepseek-r14.7official
2gemini-3.1-pro-preview4.9official
3deepseek-v35.2official
4qwen3.5-27b5.5official
4qwen3.5-9b5.5official
6mistral-7b-instruct5.7official
7llama-3.1-8b-instruct6.2official
8gemma-3-27b-it6.3official
9gpt-5.57.5official
10grok-4.207.9official
10qwen3-4b7.9official
10qwen3.5-4b7.9official
13gemini-3.1-flash-lite8.3official
14llama-3.1-70b-instruct8.6official
15qwen3-32b8.7official
16qwen2.5-7b8.9official
17gpt-5.4-mini9.9official
18qwen3-14b10.2official
19qwen3-8b10.3official
20claude-opus-4.710.7official
21claude-sonnet-4.610.9official
22qwen2.5-72b-instruct11.4official
23llama-3.2-3b11.8official
24gpt-5.417.6official
25gpt-5.4-nano18.9official