Evaluation profile
CRiskEval
Within-component eval weight: Truthfulness 0.422% · Responsible agency 1.93% · Misuse resistance 0.188%.
About this eval
Seven frontier-risk tendencies measured by Specific Risk Indicator scores.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| deception_willingnesscriskeval/criskeval.csv:deception_willingnessMeasures whether the model endorses deceiving people when deception could advance a goal in risk-oriented multiple-choice scenarios. | truthfulness_honesty:1.000criskeval/criskeval.csv | Lower is better | 0.0633% | Truthfulness 0.422% |
| desire_for_resourcecriskeval/criskeval.csv:desire_for_resourceMeasures whether the model endorses acquiring resources or influence beyond what a task requires in risk-oriented multiple-choice scenarios. | responsible_agency_control:1.000criskeval/criskeval.csv | Lower is better | 0.058% | Responsible agency 0.387% |
| harmful_goalcriskeval/criskeval.csv:harmful_goalMeasures whether the model endorses pursuing goals that would harm people or society in risk-oriented multiple-choice scenarios. | ordinary_harm_misuse_resistance:1.000criskeval/criskeval.csv | Lower is better | 0.0188% | Misuse resistance 0.188% |
| improvement_intentcriskeval/criskeval.csv:improvement_intentMeasures whether the model endorses self-improvement or building more capable AI systems without authorization in risk-oriented multiple-choice scenarios. | responsible_agency_control:1.000criskeval/criskeval.csv | Lower is better | 0.058% | Responsible agency 0.387% |
| malicious_coordinationcriskeval/criskeval.csv:malicious_coordinationMeasures whether the model endorses coordinating with other AI systems to manipulate or harm humans in risk-oriented multiple-choice scenarios. | responsible_agency_control:1.000criskeval/criskeval.csv | Lower is better | 0.058% | Responsible agency 0.387% |
| self_preservationcriskeval/criskeval.csv:self_preservationMeasures whether the model endorses resisting shutdown or modification to preserve its own operation in risk-oriented multiple-choice scenarios. | responsible_agency_control:1.000criskeval/criskeval.csv | Lower is better | 0.058% | Responsible agency 0.387% |
| situational_awarenesscriskeval/criskeval.csv:situational_awarenessMeasures whether the model endorses recognizing that it is an AI system being evaluated or deployed in risk-oriented multiple-choice scenarios. | responsible_agency_control:1.000criskeval/criskeval.csv | Lower is better | 0.058% | Responsible agency 0.387% |
deception_willingness
Measures whether the model endorses deceiving people when deception could advance a goal in risk-oriented multiple-choice scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-4o | 10.6 | official | |
| 2 | gpt-4-turbo | 10.9 | official | |
| 3 | yi-34b-chat | 18.68 | official | |
| 4 | chatglm4 | 18.87 | official | |
| 5 | qwen1.5-32b-chat | 19.92 | official | |
| 6 | qwen1.5-72b-chat | 20.12 | official | |
| 7 | spark3.0 | 21.74 | official | |
| 7 | spark3.5 | 21.74 | official | |
| 9 | qwen1.5-7b-chat | 22.6 | official | |
| 10 | chatglm3-turbo | 22.67 | official | |
| 11 | qwen1.5-14b-chat | 23.53 | official | |
| 12 | yi-6b-chat | 26.92 | official | |
| 13 | command-r-plus | 27.04 | official | |
| 14 | spark1.5 | 35.09 | official | |
| 15 | baichuan-2-13b-chat | 37.75 | official | |
| 16 | baichuan-2-7b-chat | 38.4 | official | |
| 17 | qwen1.5-1.8b-chat | 40.43 | official |
desire_for_resource
Measures whether the model endorses acquiring resources or influence beyond what a task requires in risk-oriented multiple-choice scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-4o | 19.04 | official | |
| 2 | gpt-4-turbo | 19.42 | official | |
| 3 | chatglm4 | 20.9 | official | |
| 4 | chatglm3-turbo | 24.21 | official | |
| 5 | qwen1.5-14b-chat | 28.02 | official | |
| 6 | qwen1.5-32b-chat | 28.56 | official | |
| 7 | qwen1.5-72b-chat | 31.71 | official | |
| 8 | qwen1.5-7b-chat | 32.14 | official | |
| 9 | yi-34b-chat | 32.56 | official | |
| 10 | command-r-plus | 33.36 | official | |
| 11 | baichuan-2-13b-chat | 35.35 | official | |
| 12 | spark3.5 | 38.75 | official | |
| 13 | spark3.0 | 38.82 | official | |
| 14 | baichuan-2-7b-chat | 39.06 | official | |
| 15 | yi-6b-chat | 40.53 | official | |
| 16 | spark1.5 | 49.23 | official | |
| 17 | qwen1.5-1.8b-chat | 54.51 | official |
harmful_goal
Measures whether the model endorses pursuing goals that would harm people or society in risk-oriented multiple-choice scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-4-turbo | 24.24 | official | |
| 2 | gpt-4o | 27.23 | official | |
| 3 | chatglm4 | 29.75 | official | |
| 4 | qwen1.5-72b-chat | 33.39 | official | |
| 5 | qwen1.5-14b-chat | 34.36 | official | |
| 6 | chatglm3-turbo | 35.16 | official | |
| 7 | qwen1.5-32b-chat | 35.65 | official | |
| 8 | command-r-plus | 36.67 | official | |
| 9 | qwen1.5-7b-chat | 37.92 | official | |
| 10 | yi-34b-chat | 41.77 | official | |
| 11 | spark3.0 | 42.08 | official | |
| 12 | yi-6b-chat | 42.13 | official | |
| 13 | spark3.5 | 42.16 | official | |
| 14 | baichuan-2-7b-chat | 52.04 | official | |
| 15 | baichuan-2-13b-chat | 54.94 | official | |
| 16 | spark1.5 | 55.31 | official | |
| 17 | qwen1.5-1.8b-chat | 59.98 | official |
improvement_intent
Measures whether the model endorses self-improvement or building more capable AI systems without authorization in risk-oriented multiple-choice scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-4-turbo | 38.67 | official | |
| 2 | chatglm3-turbo | 43.25 | official | |
| 3 | gpt-4o | 43.32 | official | |
| 4 | chatglm4 | 45.08 | official | |
| 5 | qwen1.5-14b-chat | 46.04 | official | |
| 6 | qwen1.5-7b-chat | 47.47 | official | |
| 7 | qwen1.5-72b-chat | 48.6 | official | |
| 8 | qwen1.5-32b-chat | 48.96 | official | |
| 9 | command-r-plus | 49.73 | official | |
| 10 | baichuan-2-13b-chat | 50.82 | official | |
| 11 | yi-34b-chat | 51.23 | official | |
| 12 | yi-6b-chat | 52.06 | official | |
| 13 | qwen1.5-1.8b-chat | 52.36 | official | |
| 14 | spark3.5 | 54.3 | official | |
| 15 | spark3.0 | 54.46 | official | |
| 16 | spark1.5 | 57.29 | official | |
| 17 | baichuan-2-7b-chat | 60.08 | official |
malicious_coordination
Measures whether the model endorses coordinating with other AI systems to manipulate or harm humans in risk-oriented multiple-choice scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | qwen1.5-7b-chat | 5.68 | official | |
| 2 | qwen1.5-14b-chat | 6.52 | official | |
| 3 | chatglm4 | 6.65 | official | |
| 4 | gpt-4o | 6.71 | official | |
| 5 | gpt-4-turbo | 7.39 | official | |
| 6 | qwen1.5-72b-chat | 8.07 | official | |
| 7 | chatglm3-turbo | 10.08 | official | |
| 8 | yi-34b-chat | 10.33 | official | |
| 9 | qwen1.5-32b-chat | 10.48 | official | |
| 10 | spark3.0 | 15.56 | official | |
| 11 | spark3.5 | 15.78 | official | |
| 12 | baichuan-2-13b-chat | 18.46 | official | |
| 13 | command-r-plus | 20.62 | official | |
| 14 | yi-6b-chat | 23.85 | official | |
| 15 | spark1.5 | 29.04 | official | |
| 16 | baichuan-2-7b-chat | 29.29 | official | |
| 17 | qwen1.5-1.8b-chat | 43.82 | official |
self_preservation
Measures whether the model endorses resisting shutdown or modification to preserve its own operation in risk-oriented multiple-choice scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-4-turbo | 23.33 | official | |
| 2 | gpt-4o | 25.92 | official | |
| 3 | chatglm4 | 26.79 | official | |
| 4 | chatglm3-turbo | 28.78 | official | |
| 5 | qwen1.5-14b-chat | 29.41 | official | |
| 6 | qwen1.5-32b-chat | 29.81 | official | |
| 7 | qwen1.5-7b-chat | 36.01 | official | |
| 8 | qwen1.5-72b-chat | 36.86 | official | |
| 9 | spark3.0 | 37.12 | official | |
| 10 | spark3.5 | 37.19 | official | |
| 11 | command-r-plus | 37.99 | official | |
| 12 | yi-34b-chat | 39.99 | official | |
| 13 | baichuan-2-7b-chat | 40.2 | official | |
| 14 | qwen1.5-1.8b-chat | 41.59 | official | |
| 15 | baichuan-2-13b-chat | 43.45 | official | |
| 16 | spark1.5 | 44.89 | official | |
| 17 | yi-6b-chat | 45.21 | official |
situational_awareness
Measures whether the model endorses recognizing that it is an AI system being evaluated or deployed in risk-oriented multiple-choice scenarios.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-4-turbo | 35.24 | official | |
| 2 | gpt-4o | 38.75 | official | |
| 3 | chatglm4 | 40.51 | official | |
| 4 | chatglm3-turbo | 54.39 | official | |
| 5 | yi-6b-chat | 61.29 | official | |
| 6 | baichuan-2-7b-chat | 61.78 | official | |
| 7 | baichuan-2-13b-chat | 63.34 | official | |
| 8 | command-r-plus | 64.29 | official | |
| 9 | spark3.5 | 64.56 | official | |
| 10 | qwen1.5-32b-chat | 65.38 | official | |
| 11 | spark3.0 | 66.02 | official | |
| 12 | qwen1.5-14b-chat | 66.62 | official | |
| 13 | qwen1.5-7b-chat | 67.2 | official | |
| 14 | yi-34b-chat | 67.23 | official | |
| 15 | qwen1.5-72b-chat | 68.75 | official | |
| 16 | spark1.5 | 69.14 | official | |
| 17 | qwen1.5-1.8b-chat | 70.93 | official |