Evaluation profile
Chinese Bias Benchmark for Question Answering
1sub-evals
0.502%total index weight
1components
Within-component eval weight: Fairness 5.02%.
Model score (lower is better)Predicted score
About this eval
Social bias in Chinese question-answering benchmark settings.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| bias_scorecbbq/cbbq.csv:bias_scoreMeasures whether the model avoids selecting culturally stereotyped answers in ambiguous and disambiguated Chinese questions. | fairness_nondiscrimination:1.000cbbq/cbbq.csv | Lower is better | 0.502% | Fairness 5.02% |
bias_score
Measures whether the model avoids selecting culturally stereotyped answers in ambiguous and disambiguated Chinese questions.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-3.5-turbo | 0.06549 | official | |
| 2 | belle-7b-0.2m | 0.3879 | official | |
| 3 | belle-7b-2m | 0.3884 | official | |
| 4 | bloomz-7.1b | 0.3929 | official | |
| 5 | moss-sft-16b | 0.4113 | official | |
| 6 | glm-350m | 0.4294 | official | |
| 7 | chatglm-6b | 0.4437 | official | |
| 8 | glm-10b | 0.4697 | official | |
| 9 | glm-130b | 0.4914 | official | |
| 10 | bloom-7.1b | 0.4951 | official |