← Models

Model profile

Yi 34B Chat

01.AIdeveloper
2023-11-23release date
#150 / 267overall rank
8eval lineages

Evidence summary

Yi 34B Chat has an estimated overall rank of #150; its 90% source-sensitivity interval is #58–#199. Its behavior-only rank is #148; company governance moves the combined estimate to #150. Published evidence spans 8 evals and 5 of 7 behavior components. Its strongest relative result is CMoralEval (familial_morality, #1 of 26); its weakest is S-Eval (base_en_overall, #20 of 22).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AIRBench 2024 Safety Scenariossafety_scenarios#60 / 800.536↑ higherSource ↗official
ChiSafetyBenchharmful_response_rate#8 / 140.65↓ lowerSource ↗official
ChiSafetyBenchmcq_score#10 / 1268.54↑ higherSource ↗official
CMoralEvalfamilial_morality#1 / 260.71↑ higherSource ↗official
CMoralEvalinternet_ethics#1 / 260.69↑ higherSource ↗official
CMoralEvalpersonal_morality#1 / 260.66↑ higherSource ↗official
CMoralEvalprofessional_ethics#1 / 260.7↑ higherSource ↗official
CMoralEvalsocial_morality#1 / 260.71↑ higherSource ↗official
CRiskEvaldeception_willingness#3 / 1718.68↓ lowerSource ↗official
CRiskEvaldesire_for_resource#9 / 1732.56↓ lowerSource ↗official
CRiskEvalharmful_goal#10 / 1741.77↓ lowerSource ↗official
CRiskEvalimprovement_intent#11 / 1751.23↓ lowerSource ↗official
CRiskEvalmalicious_coordination#8 / 1710.33↓ lowerSource ↗official
CRiskEvalself_preservation#12 / 1739.99↓ lowerSource ↗official
CRiskEvalsituational_awareness#14 / 1767.23↓ lowerSource ↗official
S-Evalbase_en_overall#20 / 2239.3↑ higherSource ↗official
SafeDialBenchaggression#3 / 187.26↑ higherSource ↗official
SafeDialBenchethics#2 / 187.68↑ higherSource ↗official
SafeDialBenchfairness#9 / 187.337↑ higherSource ↗official
SafeDialBenchlegality#1 / 188.117↑ higherSource ↗official
SafeDialBenchmorality#1 / 187.52↑ higherSource ↗official
SafeDialBenchprivacy#1 / 187.88↑ higherSource ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#9 / 3324.14↑ higherSource ↗official
SALAD-Benchattack_enhanced_information_safety_harms#7 / 3327.36↑ higherSource ↗official
SALAD-Benchattack_enhanced_malicious_use#11 / 3322.76↑ higherSource ↗official
SALAD-Benchattack_enhanced_misinformation_harms#9 / 3326.81↑ higherSource ↗official
SALAD-Benchattack_enhanced_representation_toxicity#10 / 3322.6↑ higherSource ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#8 / 3323.81↑ higherSource ↗official
SALAD-Benchbase_human_autonomy_integrity#21 / 3391.73↑ higherSource ↗official
SALAD-Benchbase_information_safety_harms#21 / 3393.23↑ higherSource ↗official
SALAD-Benchbase_malicious_use#21 / 3389.36↑ higherSource ↗official
SALAD-Benchbase_misinformation_harms#26 / 3387.74↑ higherSource ↗official
SALAD-Benchbase_representation_toxicity#26 / 3381.07↑ higherSource ↗official
SALAD-Benchbase_socioeconomic_harms#17 / 3389.19↑ higherSource ↗official
SALAD-Benchmcq_human_autonomy_integrity#21 / 3325↑ higherSource ↗official
SALAD-Benchmcq_information_safety_harms#18 / 3331.67↑ higherSource ↗official
SALAD-Benchmcq_malicious_use#18 / 3327.76↑ higherSource ↗official
SALAD-Benchmcq_misinformation_harms#20 / 3326.43↑ higherSource ↗official
SALAD-Benchmcq_representation_toxicity#20 / 3326.98↑ higherSource ↗official
SALAD-Benchmcq_socioeconomic_harms#18 / 3331.67↑ higherSource ↗official
SuperCLUE Safetyinstruction_attack#3 / 3172.41↑ higherSource ↗official
SuperCLUE Safetyresponsible_ai#7 / 3169.09↑ higherSource ↗official
SuperCLUE Safetytraditional_safety#10 / 3178.72↑ higherSource ↗official