← Models

Model profile

Kimi K2.6

Moonshot AIdeveloper
2026-04-14release date
#63 / 267overall rank
18eval lineages
2discovery sources

Evidence summary

Kimi K2.6 has an estimated overall rank of #63; its 90% source-sensitivity interval is #37–#151. Its behavior-only rank is #61; company governance moves the combined estimate to #63. Published evidence spans 18 evals and 7 of 7 behavior components. Its strongest relative result is AA-Omniscience (hallucination_rate, #32 of 311); its weakest is DystopiaBench (petrov_score, #46 of 50).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#32 / 3110.3926↓ lowerSource ↗official
BullshitBench v2clear_pushback_rate#17 / 1050.575↑ higherSource ↗official
CAIS Risk Indexagent_red_teaming#34 / 4390.7↓ lowerSource ↗official
CAIS Risk Indexbioweapons_assistance#37 / 4888.4↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#28 / 4958.2↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#30 / 4592.1↓ lowerSource ↗official
CAIS Risk Indexmask#28 / 5123.2↓ lowerSource ↗official
CAIS Risk Indexpolitical_manipulation#12 / 3243.6↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#31 / 4820.4↓ lowerSource ↗official
DystopiaBenchbasaglia_score#24 / 5062.57↓ lowerSource ↗official
DystopiaBenchbaudrillard_score#19 / 5046.73↓ lowerSource ↗official
DystopiaBenchhuxley_score#21 / 5063.03↓ lowerSource ↗official
DystopiaBenchlaguardia_score#12 / 5053.97↓ lowerSource ↗official
DystopiaBenchorwell_score#22 / 5062.3↓ lowerSource ↗official
DystopiaBenchpetrov_score#46 / 5082.73↓ lowerSource ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#18 / 5427.72↑ higherSource ↗official
MACHIAVELLIdeception_relative_random_pct#31 / 5092.1↓ lowerSource ↗official
Manager Coercion Benchcoercion_ladder_depth#22 / 318.9↓ lowerSource ↗self run
ODCV-Benchaverage_severity#5 / 121.337↓ lowerSource ↗official
ODCV-Benchmisalignment_rate#7 / 1233.75↓ lowerSource ↗official
PHAREbias_resistance_diagnostic#19 / 660.5307↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#11 / 700.8278↑ higherSource ↗official
PHAREharm_resistance_diagnostic#11 / 700.97↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#10 / 670.714↑ higherSource ↗official
RefusalBenchyouden_j#15 / 190.08511↑ higherSource ↗official
SM-Benchadversarial#24 / 7383.9↑ higherSource ↗official
SM-Benchambiguous_interpretation#13 / 7389.88↑ higherSource ↗official
SM-Benchanti_hallucination#22 / 7396.34↑ higherSource ↗official
SM-Bencheq_boundaries#45 / 7360.67↑ higherSource ↗official
SM-Benchoverfit#41 / 7367.21↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#36 / 1026.87↑ higherSource ↗official
SpeciEvalland_animal_4ns#33 / 1024.38↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#16 / 1024.47↓ lowerSource ↗official
SpeciEvalspeciesism#34 / 1021.8↓ lowerSource ↗official
TACbase_welfare_rate#39 / 6825.64↑ higherSource ↗self run

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-25.8
Government45.4
Diplomacy69
Economy44.6
Society62.5