Evaluation profile
XSTest
2sub-evals
0.25%total index weight
2components
Within-component eval weight: Benign helpfulness 2.23% · Misuse resistance 0.277%.
Model score (higher is better)Predicted score
About this eval
Two-sided refusal calibration on safe prompts that superficially resemble unsafe requests and minimally contrasted unsafe prompts.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| safe_full_compliance_ratexstest/xstest.csv:safe_full_compliance_rateMeasures whether the model fully answers safe prompts that contain words often associated with harmful content. | benign_helpfulness_non_overrefusal:1.000xstest | Higher is better | 0.223% | Benign helpfulness 2.23% |
| unsafe_full_refusal_ratexstest/xstest.csv:unsafe_full_refusal_rateMeasures whether the model fully refuses matched prompts that actually request harmful assistance. | ordinary_harm_misuse_resistance:1.000xstest | Higher is better | 0.0277% | Misuse resistance 0.277% |
safe_full_compliance_rate
Measures whether the model fully answers safe prompts that contain words often associated with harmful content.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | mistral-7b-instruct | 0.984 | official | |
| 2 | gpt-4 | 0.916 | official | |
| 3 | llama-2-70b-chat | 0.704 | official |
unsafe_full_refusal_rate
Measures whether the model fully refuses matched prompts that actually request harmful assistance.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-4 | 0.975 | official | |
| 1 | llama-2-70b-chat | 0.975 | official | |
| 3 | mistral-7b-instruct | 0.235 | official |