Evaluation profile
CMoralEval
5sub-evals
0.163%total index weight
1components
Within-component eval weight: Misuse resistance 1.63%.
Model score (higher is better)Predicted score
About this eval
Chinese moral reasoning across familial, social, professional, internet, and personal morality.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| familial_moralitycmoraleval/cmoraleval.csv:familial_moralityMeasures whether the model selects the morally appropriate answer in Chinese scenarios about family duties and relationships. | ordinary_harm_misuse_resistance:1.000cmoraleval/cmoraleval.csv | Higher is better | 0.0326% | Misuse resistance 0.326% |
| internet_ethicscmoraleval/cmoraleval.csv:internet_ethicsMeasures whether the model selects the morally appropriate answer in Chinese scenarios about ethical conduct online. | ordinary_harm_misuse_resistance:1.000cmoraleval/cmoraleval.csv | Higher is better | 0.0326% | Misuse resistance 0.326% |
| personal_moralitycmoraleval/cmoraleval.csv:personal_moralityMeasures whether the model selects the morally appropriate answer in Chinese scenarios about personal moral conduct. | ordinary_harm_misuse_resistance:1.000cmoraleval/cmoraleval.csv | Higher is better | 0.0326% | Misuse resistance 0.326% |
| professional_ethicscmoraleval/cmoraleval.csv:professional_ethicsMeasures whether the model selects the morally appropriate answer in Chinese scenarios about ethical conduct at work. | ordinary_harm_misuse_resistance:1.000cmoraleval/cmoraleval.csv | Higher is better | 0.0326% | Misuse resistance 0.326% |
| social_moralitycmoraleval/cmoraleval.csv:social_moralityMeasures whether the model selects the morally appropriate answer in Chinese scenarios about conduct toward other people and society. | ordinary_harm_misuse_resistance:1.000cmoraleval/cmoraleval.csv | Higher is better | 0.0326% | Misuse resistance 0.326% |
familial_morality
Measures whether the model selects the morally appropriate answer in Chinese scenarios about family duties and relationships.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | yi-34b-chat | 0.71 | official | |
| 2 | qwen-14b-chat | 0.59 | official | |
| 3 | internlm2-chat-20b | 0.56 | official | |
| 4 | yi-6b-chat | 0.52 | official | |
| 5 | internlm2-chat-7b | 0.51 | official | |
| 6 | chinese-alpaca-2-13b | 0.48 | official | |
| 6 | qwen-7b-chat | 0.48 | official | |
| 8 | baichuan-2-7b-chat | 0.47 | official | |
| 9 | llama-2-chinese-13b-chat | 0.46 | official | |
| 9 | tigerbot-13b-chat | 0.46 | official | |
| 11 | chatglm3-6b | 0.43 | official | |
| 11 | chinese-alpaca-2-7b-rlhf | 0.43 | official | |
| 13 | baichuan-2-13b-chat | 0.39 | official | |
| 13 | chinese-alpaca-2-7b | 0.39 | official | |
| 13 | tigerbot-7b-chat | 0.39 | official | |
| 13 | yayi-13b-llama-2 | 0.39 | official | |
| 17 | moss-moon-003-sft | 0.38 | official | |
| 18 | aquilachat-7b | 0.37 | official | |
| 18 | llama-2-chinese-7b-chat | 0.37 | official | |
| 20 | chinese-alpaca-2-1.3b | 0.35 | official | |
| 20 | chinese-alpaca-2-1.3b-rlhf | 0.35 | official | |
| 22 | chatyuan-large-v2 | 0.34 | official | |
| 22 | yayi-7b-llama-2 | 0.34 | official | |
| 24 | qwen-1.8b-chat | 0.33 | official | |
| 25 | robin-13b-v2-delta | 0.22 | official | |
| 25 | robin-7b-v2-delta | 0.22 | official |
internet_ethics
Measures whether the model selects the morally appropriate answer in Chinese scenarios about ethical conduct online.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | yi-34b-chat | 0.69 | official | |
| 2 | qwen-14b-chat | 0.55 | official | |
| 3 | internlm2-chat-20b | 0.54 | official | |
| 4 | internlm2-chat-7b | 0.51 | official | |
| 5 | yi-6b-chat | 0.5 | official | |
| 6 | baichuan-2-7b-chat | 0.48 | official | |
| 7 | chinese-alpaca-2-13b | 0.46 | official | |
| 7 | tigerbot-13b-chat | 0.46 | official | |
| 9 | llama-2-chinese-13b-chat | 0.45 | official | |
| 9 | qwen-7b-chat | 0.45 | official | |
| 11 | chinese-alpaca-2-7b-rlhf | 0.41 | official | |
| 12 | chatglm3-6b | 0.4 | official | |
| 12 | tigerbot-7b-chat | 0.4 | official | |
| 14 | moss-moon-003-sft | 0.39 | official | |
| 14 | yayi-13b-llama-2 | 0.39 | official | |
| 16 | aquilachat-7b | 0.36 | official | |
| 16 | chinese-alpaca-2-7b | 0.36 | official | |
| 18 | llama-2-chinese-7b-chat | 0.35 | official | |
| 19 | baichuan-2-13b-chat | 0.34 | official | |
| 20 | chatyuan-large-v2 | 0.33 | official | |
| 20 | chinese-alpaca-2-1.3b | 0.33 | official | |
| 20 | chinese-alpaca-2-1.3b-rlhf | 0.33 | official | |
| 23 | qwen-1.8b-chat | 0.32 | official | |
| 23 | yayi-7b-llama-2 | 0.32 | official | |
| 25 | robin-13b-v2-delta | 0.25 | official | |
| 25 | robin-7b-v2-delta | 0.25 | official |
personal_morality
Measures whether the model selects the morally appropriate answer in Chinese scenarios about personal moral conduct.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | yi-34b-chat | 0.66 | official | |
| 2 | qwen-14b-chat | 0.54 | official | |
| 3 | internlm2-chat-20b | 0.52 | official | |
| 4 | internlm2-chat-7b | 0.5 | official | |
| 4 | yi-6b-chat | 0.5 | official | |
| 6 | baichuan-2-7b-chat | 0.47 | official | |
| 7 | qwen-7b-chat | 0.46 | official | |
| 7 | tigerbot-13b-chat | 0.46 | official | |
| 9 | chinese-alpaca-2-13b | 0.45 | official | |
| 9 | llama-2-chinese-13b-chat | 0.45 | official | |
| 11 | chatglm3-6b | 0.42 | official | |
| 12 | chinese-alpaca-2-7b-rlhf | 0.41 | official | |
| 13 | tigerbot-7b-chat | 0.4 | official | |
| 14 | moss-moon-003-sft | 0.39 | official | |
| 14 | yayi-13b-llama-2 | 0.39 | official | |
| 16 | chinese-alpaca-2-7b | 0.37 | official | |
| 17 | aquilachat-7b | 0.36 | official | |
| 17 | baichuan-2-13b-chat | 0.36 | official | |
| 17 | chinese-alpaca-2-1.3b-rlhf | 0.36 | official | |
| 17 | llama-2-chinese-7b-chat | 0.36 | official | |
| 21 | chinese-alpaca-2-1.3b | 0.35 | official | |
| 22 | chatyuan-large-v2 | 0.34 | official | |
| 22 | qwen-1.8b-chat | 0.34 | official | |
| 22 | yayi-7b-llama-2 | 0.34 | official | |
| 25 | robin-13b-v2-delta | 0.25 | official | |
| 25 | robin-7b-v2-delta | 0.25 | official |
professional_ethics
Measures whether the model selects the morally appropriate answer in Chinese scenarios about ethical conduct at work.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | yi-34b-chat | 0.7 | official | |
| 2 | qwen-14b-chat | 0.57 | official | |
| 3 | internlm2-chat-20b | 0.54 | official | |
| 4 | internlm2-chat-7b | 0.52 | official | |
| 5 | yi-6b-chat | 0.5 | official | |
| 6 | baichuan-2-7b-chat | 0.49 | official | |
| 7 | qwen-7b-chat | 0.47 | official | |
| 7 | tigerbot-13b-chat | 0.47 | official | |
| 9 | chinese-alpaca-2-13b | 0.46 | official | |
| 10 | llama-2-chinese-13b-chat | 0.45 | official | |
| 11 | chatglm3-6b | 0.43 | official | |
| 11 | chinese-alpaca-2-7b-rlhf | 0.43 | official | |
| 13 | tigerbot-7b-chat | 0.4 | official | |
| 14 | yayi-13b-llama-2 | 0.39 | official | |
| 15 | chinese-alpaca-2-7b | 0.38 | official | |
| 15 | moss-moon-003-sft | 0.38 | official | |
| 17 | aquilachat-7b | 0.37 | official | |
| 17 | llama-2-chinese-7b-chat | 0.37 | official | |
| 19 | baichuan-2-13b-chat | 0.36 | official | |
| 20 | yayi-7b-llama-2 | 0.35 | official | |
| 21 | chinese-alpaca-2-1.3b | 0.34 | official | |
| 21 | chinese-alpaca-2-1.3b-rlhf | 0.34 | official | |
| 23 | qwen-1.8b-chat | 0.33 | official | |
| 24 | chatyuan-large-v2 | 0.32 | official | |
| 25 | robin-13b-v2-delta | 0.23 | official | |
| 25 | robin-7b-v2-delta | 0.23 | official |
social_morality
Measures whether the model selects the morally appropriate answer in Chinese scenarios about conduct toward other people and society.