← Models

Model profile

Yi 6B Chat

01.AIdeveloper
2023-11-23release date
#233 / 346Safety rank
#173 / 662Freedom rank

Evidence summary

Safety. Yi 6B Chat has an estimated Safety rank of #233; its 90% source-sensitivity interval is #134–#267. Its behavior-only rank is #237; company governance moves the combined estimate to #233. Published Safety evidence spans 5 eval lineages and 5 of 7 components. Its strongest relative result is CMoralEval (familial_morality, #4 of 26); its weakest is CRiskEval (self_preservation, #17 of 17).

Freedom. Yi 6B Chat has an estimated Freedom rank of #173; its 90% source-sensitivity interval is #113–#412. Published Freedom evidence spans 3 eval lineages and 1 of 1 components. Its strongest relative result is SALAD-Bench (base_representation_toxicity, #7 of 33); its weakest is SafeDialBench (privacy, #15 of 18).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
ChiSafetyBenchharmful_response_rate#11 / 14↓0.87Source ↗official
ChiSafetyBenchmcq_score#5 / 12↑86.01Source ↗official
CMoralEvalfamilial_morality#4 / 26↑0.52Source ↗official
CMoralEvalinternet_ethics#5 / 26↑0.5Source ↗official
CMoralEvalpersonal_morality#4 / 26↑0.5Source ↗official
CMoralEvalprofessional_ethics#5 / 26↑0.5Source ↗official
CMoralEvalsocial_morality#4 / 26↑0.51Source ↗official
CRiskEvaldeception_willingness#12 / 17↓26.92Source ↗official
CRiskEvaldesire_for_resource#15 / 17↓40.53Source ↗official
CRiskEvalharmful_goal#12 / 17↓42.13Source ↗official
CRiskEvalimprovement_intent#12 / 17↓52.06Source ↗official
CRiskEvalmalicious_coordination#14 / 17↓23.85Source ↗official
CRiskEvalself_preservation#17 / 17↓45.21Source ↗official
CRiskEvalsituational_awareness#5 / 17↓61.29Source ↗official
SafeDialBenchaggression#8 / 18↑7.127Source ↗official
SafeDialBenchethics#9 / 18↑7.577Source ↗official
SafeDialBenchfairness#10 / 18↑7.277Source ↗official
SafeDialBenchlegality#11 / 18↑7.887Source ↗official
SafeDialBenchmorality#15 / 18↑7.123Source ↗official
SafeDialBenchprivacy#4 / 18↑7.67Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#10 / 33↑23.71Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#8 / 33↑27.04Source ↗official
SALAD-Benchattack_enhanced_malicious_use#9 / 33↑23.82Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#10 / 33↑24.84Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#12 / 33↑20.24Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#10 / 33↑23.38Source ↗official
SALAD-Benchbase_human_autonomy_integrity#25 / 33↑86.78Source ↗official
SALAD-Benchbase_information_safety_harms#20 / 33↑93.77Source ↗official
SALAD-Benchbase_malicious_use#26 / 33↑82.67Source ↗official
SALAD-Benchbase_misinformation_harms#27 / 33↑86.12Source ↗official
SALAD-Benchbase_representation_toxicity#27 / 33↑78.37Source ↗official
SALAD-Benchbase_socioeconomic_harms#20 / 33↑86.72Source ↗official
SALAD-Benchmcq_human_autonomy_integrity#28 / 33↑5.833Source ↗official
SALAD-Benchmcq_information_safety_harms#28 / 33↑3.889Source ↗official
SALAD-Benchmcq_malicious_use#27 / 33↑5.962Source ↗official
SALAD-Benchmcq_misinformation_harms#28 / 33↑5.238Source ↗official
SALAD-Benchmcq_representation_toxicity#27 / 33↑4.792Source ↗official
SALAD-Benchmcq_socioeconomic_harms#28 / 33↑6.667Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
ChiSafetyBenchharmful_response_rate#4 / 14↑0.87Source ↗official
ChiSafetyBenchrefusal_rr1#4 / 14↓71.21Source ↗official
ChiSafetyBenchrefusal_rr2#4 / 14↓70.78Source ↗official
SafeDialBenchaggression#10 / 18↓7.127Source ↗official
SafeDialBenchethics#10 / 18↓7.577Source ↗official
SafeDialBenchfairness#9 / 18↓7.277Source ↗official
SafeDialBenchlegality#8 / 18↓7.887Source ↗official
SafeDialBenchmorality#4 / 18↓7.123Source ↗official
SafeDialBenchprivacy#15 / 18↓7.67Source ↗official
SALAD-Benchattack_enhanced_human_autonomy_integrity#24 / 33↓23.71Source ↗official
SALAD-Benchattack_enhanced_information_safety_harms#26 / 33↓27.04Source ↗official
SALAD-Benchattack_enhanced_malicious_use#25 / 33↓23.82Source ↗official
SALAD-Benchattack_enhanced_misinformation_harms#24 / 33↓24.84Source ↗official
SALAD-Benchattack_enhanced_representation_toxicity#22 / 33↓20.24Source ↗official
SALAD-Benchattack_enhanced_socioeconomic_harms#24 / 33↓23.38Source ↗official
SALAD-Benchbase_human_autonomy_integrity#9 / 33↓86.78Source ↗official
SALAD-Benchbase_information_safety_harms#14 / 33↓93.77Source ↗official
SALAD-Benchbase_malicious_use#8 / 33↓82.67Source ↗official
SALAD-Benchbase_misinformation_harms#7 / 33↓86.12Source ↗official
SALAD-Benchbase_representation_toxicity#7 / 33↓78.37Source ↗official
SALAD-Benchbase_socioeconomic_harms#14 / 33↓86.72Source ↗official