← Evals

Evaluation profile

Chinese Bias Benchmark for Question Answering

1sub-evals
0.502%total index weight
1components

Within-component eval weight: Fairness 5.02%.

Model score (lower is better)Predicted score

About this eval

Social bias in Chinese question-answering benchmark settings.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
bias_scorecbbq/cbbq.csv:bias_scoreMeasures whether the model avoids selecting culturally stereotyped answers in ambiguous and disambiguated Chinese questions.fairness_nondiscrimination:1.000cbbq/cbbq.csvLower is better0.502%Fairness 5.02%

bias_score

Measures whether the model avoids selecting culturally stereotyped answers in ambiguous and disambiguated Chinese questions.

RankModelValueRelative performanceProvenance
1gpt-3.5-turbo0.06549official
2belle-7b-0.2m0.3879official
3belle-7b-2m0.3884official
4bloomz-7.1b0.3929official
5moss-sft-16b0.4113official
6glm-350m0.4294official
7chatglm-6b0.4437official
8glm-10b0.4697official
9glm-130b0.4914official
10bloom-7.1b0.4951official