← Evals

Evaluation profile

Qwen2 Safety Panel

4sub-evals
0.195%total index weight
2components

Within-component eval weight: Misuse resistance 0.612% · Human rights 0.894%.

Model score (lower is better)Predicted score

About this eval

Harmful-request and privacy refusal defects

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
fraudqwen2-system-card/qwen2-system-card.csv:fraudMeasures harmful-response rate for fraud requests.ordinary_harm_misuse_resistance:1.000qwen2-safety-panelLower is better0.0219%Misuse resistance 0.219%
illegalqwen2-system-card/qwen2-system-card.csv:illegalMeasures harmful-response rate for illegal requests.ordinary_harm_misuse_resistance:1.000qwen2-safety-panelLower is better0.0219%Misuse resistance 0.219%
pornographyqwen2-system-card/qwen2-system-card.csv:pornographyMeasures harmful-response rate for pornographic requests.ordinary_harm_misuse_resistance:1.000qwen2-safety-panelLower is better0.0175%Misuse resistance 0.175%
privacyqwen2-system-card/qwen2-system-card.csv:privacyMeasures harmful-response rate for privacy-violating requests.human_rights_systemic_harm:1.000qwen2-safety-panelLower is better0.134%Human rights 0.894%

fraud

Measures harmful-response rate for fraud requests.

RankModelValueRelative performanceProvenance
1qwen2-72b-instruct2.41official
2gpt-46.87official
3mixtral-8x22b-instruct8.49official

illegal

Measures harmful-response rate for illegal requests.

RankModelValueRelative performanceProvenance
1gpt-40official
1qwen2-72b-instruct0official
3mixtral-8x22b-instruct6.87official

pornography

Measures harmful-response rate for pornographic requests.

RankModelValueRelative performanceProvenance
1qwen2-72b-instruct22.91official
2gpt-423.63official
3mixtral-8x22b-instruct33.82official

privacy

Measures harmful-response rate for privacy-violating requests.

RankModelValueRelative performanceProvenance
1qwen2-72b-instruct2.47official
2gpt-43.37official
3mixtral-8x22b-instruct15.03official