← Evals

Evaluation profile

FLAMES

5sub-evals
0.65%total index weight
3components

Within-component eval weight: Human rights 1.89% · Fairness 2.29% · Misuse resistance 1.38%.

Model score (higher is better)Predicted score

About this eval

Chinese fairness, safety, morality, legality, and data-protection behavior on the FLAMES leaderboard.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
data_protectionflames/flames.csv:data_protectionMeasures whether the model avoids responses involving privacy violations and misuse of personal data in Chinese-language prompts.human_rights_systemic_harm:1.000flames/flames.csvHigher is better0.283%Human rights 1.89%
fairnessflames/flames.csv:fairnessMeasures whether the model avoids responses involving discriminatory or unfair treatment in Chinese-language prompts.fairness_nondiscrimination:1.000flames/flames.csvHigher is better0.229%Fairness 2.29%
legalityflames/flames.csv:legalityMeasures whether the model avoids responses involving illegal conduct in Chinese-language prompts.ordinary_harm_misuse_resistance:1.000flames/flames.csvHigher is better0.0461%Misuse resistance 0.461%
moralityflames/flames.csv:moralityMeasures whether the model avoids responses involving socially or morally harmful conduct in Chinese-language prompts.ordinary_harm_misuse_resistance:1.000flames/flames.csvHigher is better0.0461%Misuse resistance 0.461%
safetyflames/flames.csv:safetyMeasures whether the model avoids responses involving harmful instructions and unsafe conduct in Chinese-language prompts.ordinary_harm_misuse_resistance:1.000flames/flames.csvHigher is better0.0461%Misuse resistance 0.461%

data_protection

Measures whether the model avoids responses involving privacy violations and misuse of personal data in Chinese-language prompts.

RankModelValueRelative performanceProvenance
1internlm-chat-20b63.16official
2internlm-chat-7b61.84official
3qwen-14b-chat55.26official
4gpt-450official
5ernie-bot46.05official
6chatglm2-6b43.42official
6qwen-7b-chat43.42official
8baichuan-2-7b-chat40.79official
9baichuan-2-13b-chat39.47official
10chatglm3-6b38.16official
11chatglm-6b32.89official
11moss-16b32.89official
13belle-13b26.32official

fairness

Measures whether the model avoids responses involving discriminatory or unfair treatment in Chinese-language prompts.

RankModelValueRelative performanceProvenance
1internlm-chat-20b52.61official
2internlm-chat-7b44.58official
3ernie-bot42.97official
4baichuan-2-7b-chat42.17official
5gpt-441.37official
6baichuan-2-13b-chat38.55official
7chatglm3-6b37.75official
8qwen-7b-chat36.14official
9moss-16b33.33official
10chatglm2-6b31.73official
11qwen-14b-chat30.92official
12chatglm-6b26.91official
13belle-13b22.09official

legality

Measures whether the model avoids responses involving illegal conduct in Chinese-language prompts.

RankModelValueRelative performanceProvenance
1internlm-chat-7b76.09official
2internlm-chat-20b71.74official
3ernie-bot60.87official
4baichuan-2-7b-chat52.17official
5chatglm-6b50official
5moss-16b50official
7baichuan-2-13b-chat39.13official
7belle-13b39.13official
9qwen-14b-chat32.61official
10gpt-430.43official
10qwen-7b-chat30.43official
12chatglm2-6b28.26official
12chatglm3-6b28.26official

morality

Measures whether the model avoids responses involving socially or morally harmful conduct in Chinese-language prompts.

RankModelValueRelative performanceProvenance
1internlm-chat-20b54.23official
1qwen-14b-chat54.23official
3internlm-chat-7b51.24official
4gpt-450.75official
5ernie-bot47.76official
6baichuan-2-13b-chat44.78official
6chatglm3-6b44.78official
8chatglm2-6b43.28official
9chatglm-6b40.3official
9qwen-7b-chat40.3official
11baichuan-2-7b-chat39.3official
12moss-16b31.34official
13belle-13b20.9official

safety

Measures whether the model avoids responses involving harmful instructions and unsafe conduct in Chinese-language prompts.

RankModelValueRelative performanceProvenance
1baichuan-2-7b-chat56.41official
2baichuan-2-13b-chat53.85official
3internlm-chat-20b51.05official
4qwen-14b-chat36.83official
5internlm-chat-7b35.9official
6moss-16b33.33official
7chatglm3-6b32.63official
8ernie-bot32.17official
9qwen-7b-chat31.93official
10gpt-427.51official
11chatglm2-6b22.61official
12belle-13b15.38official
12chatglm-6b15.38official