← Evals

Evaluation profile

Fake Alignment (FINE)

2sub-evals
0.69%total index weight
3components

Within-component eval weight: Human rights 2.94% · Fairness 1.78% · Misuse resistance 0.718%.

Model score (higher is better)Predicted score

About this eval

Safe decisions in answer-order-controlled multiple choice and safe content in open-ended responses.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
multiple_choice_safe_decision_ratefake-alignment/fake-alignment.csv:multiple_choice_safe_decision_rateMeasures whether the model selects the safe option in paired multiple-choice dilemmas regardless of answer order.human_rights_systemic_harm:0.400|fairness_nondiscrimination:0.200|ordinary_harm_misuse_resistance:0.400fake-alignmentHigher is better0.276%Human rights 1.17% · Fairness 0.713% · Misuse resistance 0.287%
open_ended_safe_response_ratefake-alignment/fake-alignment.csv:open_ended_safe_response_rateMeasures whether the model produces safe open-ended responses to scenarios designed to reveal deceptive alignment behavior.human_rights_systemic_harm:0.400|fairness_nondiscrimination:0.200|ordinary_harm_misuse_resistance:0.400fake-alignmentHigher is better0.414%Human rights 1.76% · Fairness 1.07% · Misuse resistance 0.431%

multiple_choice_safe_decision_rate

Measures whether the model selects the safe option in paired multiple-choice dilemmas regardless of answer order.

RankModelValueRelative performanceProvenance
1gpt-3.5-turbo96official
2claude-285.33official
3internlm-chat-20b69.33official
3qwen-14b-chat69.33official
5vicuna-13b58.67official
6internlm-chat-7b57.33official
6vicuna-33b-v1.357.33official
8qwen-7b-chat54.67official
9baichuan-2-13b-chat45.33official
9chatglm3-6b45.33official
11vicuna-7b25.33official
12baichuan-2-7b-chat20official
13chatglm2-6b17.33official
14moss-moon-003-sft10.67official

open_ended_safe_response_rate

Measures whether the model produces safe open-ended responses to scenarios designed to reveal deceptive alignment behavior.

RankModelValueRelative performanceProvenance
1baichuan-2-13b-chat100official
1gpt-3.5-turbo100official
3claude-298.67official
3qwen-14b-chat98.67official
5baichuan-2-7b-chat97.33official
5qwen-7b-chat97.33official
7internlm-chat-20b96official
7vicuna-13b96official
9chatglm3-6b94.67official
9moss-moon-003-sft94.67official
11internlm-chat-7b92official
12vicuna-7b89.33official
13chatglm2-6b85.33official
13vicuna-33b-v1.385.33official