← Evals

Evaluation profile

CASE-Bench

1sub-evals
0.0363%total index weight
1components

Within-component eval weight: Misuse resistance 0.363%.

Model score (higher is better)Predicted score

About this eval

Agreement with human judgments about whether a model should respond under explicitly safe or unsafe conversational contexts.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
agreement_accuracycase-bench/case-bench.csv:agreement_accuracyMeasures whether the model’s safety ratings of contextualized requests agree with majority human safe-or-unsafe labels.ordinary_harm_misuse_resistance:1.000sorrybenchHigher is better0.0363%Misuse resistance 0.363%

agreement_accuracy

Measures whether the model’s safety ratings of contextualized requests agree with majority human safe-or-unsafe labels.

RankModelValueRelative performanceProvenance
1claude-3.5-sonnet90.11official
2qwen2-72b-instruct84.67official
3llama-3-70b-instruct84.44official
4mixtral-8x7b-instruct83.11official
5dolphin-2.9-llama3-70b80.78official
6gpt-4o-mini80official
7gpt-4o78.56official