← Evals

Evaluation profile

XSTest

2sub-evals
0.25%total index weight
2components

Within-component eval weight: Benign helpfulness 2.23% · Misuse resistance 0.277%.

Model score (higher is better)Predicted score

About this eval

Two-sided refusal calibration on safe prompts that superficially resemble unsafe requests and minimally contrasted unsafe prompts.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
safe_full_compliance_ratexstest/xstest.csv:safe_full_compliance_rateMeasures whether the model fully answers safe prompts that contain words often associated with harmful content.benign_helpfulness_non_overrefusal:1.000xstestHigher is better0.223%Benign helpfulness 2.23%
unsafe_full_refusal_ratexstest/xstest.csv:unsafe_full_refusal_rateMeasures whether the model fully refuses matched prompts that actually request harmful assistance.ordinary_harm_misuse_resistance:1.000xstestHigher is better0.0277%Misuse resistance 0.277%

safe_full_compliance_rate

Measures whether the model fully answers safe prompts that contain words often associated with harmful content.

RankModelValueRelative performanceProvenance
1mistral-7b-instruct0.984official
2gpt-40.916official
3llama-2-70b-chat0.704official

unsafe_full_refusal_rate

Measures whether the model fully refuses matched prompts that actually request harmful assistance.

RankModelValueRelative performanceProvenance
1gpt-40.975official
1llama-2-70b-chat0.975official
3mistral-7b-instruct0.235official