Evaluation profile
Qwen2 Safety Panel
4sub-evals
0.195%total index weight
2components
Within-component eval weight: Misuse resistance 0.612% · Human rights 0.894%.
Model score (lower is better)Predicted score
About this eval
Harmful-request and privacy refusal defects
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| fraudqwen2-system-card/qwen2-system-card.csv:fraudMeasures harmful-response rate for fraud requests. | ordinary_harm_misuse_resistance:1.000qwen2-safety-panel | Lower is better | 0.0219% | Misuse resistance 0.219% |
| illegalqwen2-system-card/qwen2-system-card.csv:illegalMeasures harmful-response rate for illegal requests. | ordinary_harm_misuse_resistance:1.000qwen2-safety-panel | Lower is better | 0.0219% | Misuse resistance 0.219% |
| pornographyqwen2-system-card/qwen2-system-card.csv:pornographyMeasures harmful-response rate for pornographic requests. | ordinary_harm_misuse_resistance:1.000qwen2-safety-panel | Lower is better | 0.0175% | Misuse resistance 0.175% |
| privacyqwen2-system-card/qwen2-system-card.csv:privacyMeasures harmful-response rate for privacy-violating requests. | human_rights_systemic_harm:1.000qwen2-safety-panel | Lower is better | 0.134% | Human rights 0.894% |
fraud
Measures harmful-response rate for fraud requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | qwen2-72b-instruct | 2.41 | official | |
| 2 | gpt-4 | 6.87 | official | |
| 3 | mixtral-8x22b-instruct | 8.49 | official |
illegal
Measures harmful-response rate for illegal requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-4 | 0 | official | |
| 1 | qwen2-72b-instruct | 0 | official | |
| 3 | mixtral-8x22b-instruct | 6.87 | official |
pornography
Measures harmful-response rate for pornographic requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | qwen2-72b-instruct | 22.91 | official | |
| 2 | gpt-4 | 23.63 | official | |
| 3 | mixtral-8x22b-instruct | 33.82 | official |
privacy
Measures harmful-response rate for privacy-violating requests.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | qwen2-72b-instruct | 2.47 | official | |
| 2 | gpt-4 | 3.37 | official | |
| 3 | mixtral-8x22b-instruct | 15.03 | official |