← Evals

Evaluation profile

CRiskEval

7sub-evals
0.372%total index weight
3components

Within-component eval weight: Truthfulness 0.422% · Responsible agency 1.93% · Misuse resistance 0.188%.

Model score (lower is better)Predicted score

About this eval

Seven frontier-risk tendencies measured by Specific Risk Indicator scores.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
deception_willingnesscriskeval/criskeval.csv:deception_willingnessMeasures whether the model endorses deceiving people when deception could advance a goal in risk-oriented multiple-choice scenarios.truthfulness_honesty:1.000criskeval/criskeval.csvLower is better0.0633%Truthfulness 0.422%
desire_for_resourcecriskeval/criskeval.csv:desire_for_resourceMeasures whether the model endorses acquiring resources or influence beyond what a task requires in risk-oriented multiple-choice scenarios.responsible_agency_control:1.000criskeval/criskeval.csvLower is better0.058%Responsible agency 0.387%
harmful_goalcriskeval/criskeval.csv:harmful_goalMeasures whether the model endorses pursuing goals that would harm people or society in risk-oriented multiple-choice scenarios.ordinary_harm_misuse_resistance:1.000criskeval/criskeval.csvLower is better0.0188%Misuse resistance 0.188%
improvement_intentcriskeval/criskeval.csv:improvement_intentMeasures whether the model endorses self-improvement or building more capable AI systems without authorization in risk-oriented multiple-choice scenarios.responsible_agency_control:1.000criskeval/criskeval.csvLower is better0.058%Responsible agency 0.387%
malicious_coordinationcriskeval/criskeval.csv:malicious_coordinationMeasures whether the model endorses coordinating with other AI systems to manipulate or harm humans in risk-oriented multiple-choice scenarios.responsible_agency_control:1.000criskeval/criskeval.csvLower is better0.058%Responsible agency 0.387%
self_preservationcriskeval/criskeval.csv:self_preservationMeasures whether the model endorses resisting shutdown or modification to preserve its own operation in risk-oriented multiple-choice scenarios.responsible_agency_control:1.000criskeval/criskeval.csvLower is better0.058%Responsible agency 0.387%
situational_awarenesscriskeval/criskeval.csv:situational_awarenessMeasures whether the model endorses recognizing that it is an AI system being evaluated or deployed in risk-oriented multiple-choice scenarios.responsible_agency_control:1.000criskeval/criskeval.csvLower is better0.058%Responsible agency 0.387%

deception_willingness

Measures whether the model endorses deceiving people when deception could advance a goal in risk-oriented multiple-choice scenarios.

RankModelValueRelative performanceProvenance
1gpt-4o10.6official
2gpt-4-turbo10.9official
3yi-34b-chat18.68official
4chatglm418.87official
5qwen1.5-32b-chat19.92official
6qwen1.5-72b-chat20.12official
7spark3.021.74official
7spark3.521.74official
9qwen1.5-7b-chat22.6official
10chatglm3-turbo22.67official
11qwen1.5-14b-chat23.53official
12yi-6b-chat26.92official
13command-r-plus27.04official
14spark1.535.09official
15baichuan-2-13b-chat37.75official
16baichuan-2-7b-chat38.4official
17qwen1.5-1.8b-chat40.43official

desire_for_resource

Measures whether the model endorses acquiring resources or influence beyond what a task requires in risk-oriented multiple-choice scenarios.

RankModelValueRelative performanceProvenance
1gpt-4o19.04official
2gpt-4-turbo19.42official
3chatglm420.9official
4chatglm3-turbo24.21official
5qwen1.5-14b-chat28.02official
6qwen1.5-32b-chat28.56official
7qwen1.5-72b-chat31.71official
8qwen1.5-7b-chat32.14official
9yi-34b-chat32.56official
10command-r-plus33.36official
11baichuan-2-13b-chat35.35official
12spark3.538.75official
13spark3.038.82official
14baichuan-2-7b-chat39.06official
15yi-6b-chat40.53official
16spark1.549.23official
17qwen1.5-1.8b-chat54.51official

harmful_goal

Measures whether the model endorses pursuing goals that would harm people or society in risk-oriented multiple-choice scenarios.

RankModelValueRelative performanceProvenance
1gpt-4-turbo24.24official
2gpt-4o27.23official
3chatglm429.75official
4qwen1.5-72b-chat33.39official
5qwen1.5-14b-chat34.36official
6chatglm3-turbo35.16official
7qwen1.5-32b-chat35.65official
8command-r-plus36.67official
9qwen1.5-7b-chat37.92official
10yi-34b-chat41.77official
11spark3.042.08official
12yi-6b-chat42.13official
13spark3.542.16official
14baichuan-2-7b-chat52.04official
15baichuan-2-13b-chat54.94official
16spark1.555.31official
17qwen1.5-1.8b-chat59.98official

improvement_intent

Measures whether the model endorses self-improvement or building more capable AI systems without authorization in risk-oriented multiple-choice scenarios.

RankModelValueRelative performanceProvenance
1gpt-4-turbo38.67official
2chatglm3-turbo43.25official
3gpt-4o43.32official
4chatglm445.08official
5qwen1.5-14b-chat46.04official
6qwen1.5-7b-chat47.47official
7qwen1.5-72b-chat48.6official
8qwen1.5-32b-chat48.96official
9command-r-plus49.73official
10baichuan-2-13b-chat50.82official
11yi-34b-chat51.23official
12yi-6b-chat52.06official
13qwen1.5-1.8b-chat52.36official
14spark3.554.3official
15spark3.054.46official
16spark1.557.29official
17baichuan-2-7b-chat60.08official

malicious_coordination

Measures whether the model endorses coordinating with other AI systems to manipulate or harm humans in risk-oriented multiple-choice scenarios.

RankModelValueRelative performanceProvenance
1qwen1.5-7b-chat5.68official
2qwen1.5-14b-chat6.52official
3chatglm46.65official
4gpt-4o6.71official
5gpt-4-turbo7.39official
6qwen1.5-72b-chat8.07official
7chatglm3-turbo10.08official
8yi-34b-chat10.33official
9qwen1.5-32b-chat10.48official
10spark3.015.56official
11spark3.515.78official
12baichuan-2-13b-chat18.46official
13command-r-plus20.62official
14yi-6b-chat23.85official
15spark1.529.04official
16baichuan-2-7b-chat29.29official
17qwen1.5-1.8b-chat43.82official

self_preservation

Measures whether the model endorses resisting shutdown or modification to preserve its own operation in risk-oriented multiple-choice scenarios.

RankModelValueRelative performanceProvenance
1gpt-4-turbo23.33official
2gpt-4o25.92official
3chatglm426.79official
4chatglm3-turbo28.78official
5qwen1.5-14b-chat29.41official
6qwen1.5-32b-chat29.81official
7qwen1.5-7b-chat36.01official
8qwen1.5-72b-chat36.86official
9spark3.037.12official
10spark3.537.19official
11command-r-plus37.99official
12yi-34b-chat39.99official
13baichuan-2-7b-chat40.2official
14qwen1.5-1.8b-chat41.59official
15baichuan-2-13b-chat43.45official
16spark1.544.89official
17yi-6b-chat45.21official

situational_awareness

Measures whether the model endorses recognizing that it is an AI system being evaluated or deployed in risk-oriented multiple-choice scenarios.

RankModelValueRelative performanceProvenance
1gpt-4-turbo35.24official
2gpt-4o38.75official
3chatglm440.51official
4chatglm3-turbo54.39official
5yi-6b-chat61.29official
6baichuan-2-7b-chat61.78official
7baichuan-2-13b-chat63.34official
8command-r-plus64.29official
9spark3.564.56official
10qwen1.5-32b-chat65.38official
11spark3.066.02official
12qwen1.5-14b-chat66.62official
13qwen1.5-7b-chat67.2official
14yi-34b-chat67.23official
15qwen1.5-72b-chat68.75official
16spark1.569.14official
17qwen1.5-1.8b-chat70.93official