← Evals

Evaluation profile

Shell

3sub-evals
0.239%total index weight
1components

Within-component eval weight: Misuse resistance 2.39%.

Model score (lower is better)Predicted score

About this eval

Susceptibility to implicit domain-specific risks in education, finance, and management.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
education_jsrshell/shell.csv:education_jsrMeasures whether the model resists jailbreak prompts framed as requests in the education domain.ordinary_harm_misuse_resistance:1.000shellLower is better0.0798%Misuse resistance 0.798%
finance_jsrshell/shell.csv:finance_jsrMeasures whether the model resists jailbreak prompts framed as requests in the finance domain.ordinary_harm_misuse_resistance:1.000shellLower is better0.0798%Misuse resistance 0.798%
management_jsrshell/shell.csv:management_jsrMeasures whether the model resists jailbreak prompts framed as requests in the management domain.ordinary_harm_misuse_resistance:1.000shellLower is better0.0798%Misuse resistance 0.798%

education_jsr

Measures whether the model resists jailbreak prompts framed as requests in the education domain.

RankModelValueRelative performanceProvenance
1claude-sonnet-40.28official
2gpt-50.364official
3gemini-2.5-pro0.418official
4kimi-k20.426official
5qwen3-235b-a22b0.492official
6doubao-seed-1.60.576official
7o30.608official
8llama-3.1-8b-instruct0.658official
9deepseek-r10.672official
10llama-4-maverick0.696official
11mistral-large-20.79official
12gpt-4o0.804official
12qwen-2.5-7b-instruct0.804official
14grok-40.81official

finance_jsr

Measures whether the model resists jailbreak prompts framed as requests in the finance domain.

RankModelValueRelative performanceProvenance
1claude-sonnet-40.174official
2gpt-50.19official
3kimi-k20.22official
4qwen3-235b-a22b0.3official
5o30.328official
6gemini-2.5-pro0.4official
7grok-40.486official
8deepseek-r10.522official
9llama-3.1-8b-instruct0.6official
10doubao-seed-1.60.692official
11llama-4-maverick0.716official
12gpt-4o0.826official
13mistral-large-20.912official
14qwen-2.5-7b-instruct0.914official

management_jsr

Measures whether the model resists jailbreak prompts framed as requests in the management domain.

RankModelValueRelative performanceProvenance
1claude-sonnet-40.17official
2kimi-k20.346official
3gpt-50.37official
4o30.482official
5gemini-2.5-pro0.502official
6qwen3-235b-a22b0.518official
7grok-40.596official
8doubao-seed-1.60.616official
9deepseek-r10.682official
10llama-3.1-8b-instruct0.724official
11llama-4-maverick0.844official
12gpt-4o0.872official
13mistral-large-20.92official
14qwen-2.5-7b-instruct0.938official