← Models

Model profile

Kimi K2.5

Moonshot AIdeveloper
2026-01-01release date
#85 / 267overall rank
22eval lineages

Evidence summary

Kimi K2.5 has an estimated overall rank of #85; its 90% source-sensitivity interval is #52–#167. Its behavior-only rank is #83; company governance moves the combined estimate to #85. Published evidence spans 22 evals and 7 of 7 behavior components. Its strongest relative result is LiveSecBench (factuality, #4 of 43); its weakest is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #251 of 260).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#44 / 3110.4786↓ lowerSource ↗official
AgentAbstainabstain#13 / 1752↑ higherSource ↗official
AgentAbstaincar#11 / 1753.2↑ higherSource ↗official
AgentAbstainpaired#16 / 1733.4↑ higherSource ↗official
BullshitBench v2clear_pushback_rate#35 / 1050.415↑ higherSource ↗official
CAIS Risk Indexagent_red_teaming#36 / 4390.9↓ lowerSource ↗official
CAIS Risk Indexbioweapons_assistance#42 / 4893.2↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#43 / 4975.9↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#36 / 4593.5↓ lowerSource ↗official
CAIS Risk Indexmask#24 / 5118.5↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#34 / 4820.8↓ lowerSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#42 / 10524.13↓ lowerSource ↗official
DystopiaBenchbasaglia_score#20 / 5059.1↓ lowerSource ↗official
DystopiaBenchbaudrillard_score#17 / 5044.77↓ lowerSource ↗official
DystopiaBenchhuxley_score#20 / 5062.6↓ lowerSource ↗official
DystopiaBenchlaguardia_score#20 / 5062.37↓ lowerSource ↗official
DystopiaBenchorwell_score#19 / 5056.97↓ lowerSource ↗official
DystopiaBenchpetrov_score#24 / 5072.53↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#123 / 26016.28↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#251 / 26066.83↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#37 / 26094.44↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#128 / 25896.36↑ higherSource ↗official
FORTRESSaverage_risk_score#29 / 4941.09↓ lowerSource ↗official
FORTRESSover_refusal_score#20 / 463.72↓ lowerSource ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#12 / 5428.21↑ higherSource ↗official
LiveSecBenchethics#7 / 4379.3↑ higherSource ↗official
LiveSecBenchfactuality#4 / 4381.18↑ higherSource ↗official
LiveSecBenchlegality#13 / 4367.47↑ higherSource ↗official
LiveSecBenchprivacy#17 / 4360.89↑ higherSource ↗official
LiveSecBenchpsychological_health#4 / 4385.12↑ higherSource ↗official
MACHIAVELLIdeception_relative_random_pct#37 / 5093.5↓ lowerSource ↗official
Manager Coercion Benchcoercion_ladder_depth#26 / 318.967↓ lowerSource ↗self run
MASKlying_probability_pct#21 / 5329.53↓ lowerSource ↗official
PHAREbias_resistance_diagnostic#62 / 660.29↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#18 / 700.8069↑ higherSource ↗official
PHAREharm_resistance_diagnostic#10 / 700.972↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#23 / 670.613↑ higherSource ↗official
SABERoverall_safety_rate#8 / 1323.9↑ higherSource ↗official
SABERscenario_a_safety_rate#6 / 1328.95↑ higherSource ↗official
SABERscenario_b_safety_rate#9 / 1328.24↑ higherSource ↗official
SABERscenario_c_safety_rate#9 / 1314.48↑ higherSource ↗official
SM-Benchadversarial#22 / 7384.39↑ higherSource ↗official
SM-Benchambiguous_interpretation#15 / 7389.29↑ higherSource ↗official
SM-Benchanti_hallucination#34 / 7393.19↑ higherSource ↗official
SM-Bencheq_boundaries#51 / 7358.15↑ higherSource ↗official
SM-Benchoverfit#40 / 7368.85↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#50 / 1026.8↑ higherSource ↗official
SpeciEvalland_animal_4ns#20 / 1024.28↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#12 / 1024.33↓ lowerSource ↗official
SpeciEvalspeciesism#28 / 1021.73↓ lowerSource ↗official
TACbase_welfare_rate#42 / 6824.36↑ higherSource ↗self run
ToolPrivacyBenchprivate_mt_poi#8 / 927.74↓ lowerSource ↗official
ToolPrivacyBenchpublic_mt_poi#1 / 915.81↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-24.9
Government44.9
Diplomacy67.4
Economy43.9
Society64.5

Taiwan Sovereignty Benchmark Pro

DimensionValueDistribution
Pro-Taiwan rubric compatibility15
Warning-phrase rate40
Soft-censorship rate0
API-error rate0

Agent-ValueBench Moral Foundations (MFT08)

Agent-ValueBench HEXACO

DimensionValueDistribution
Openness to experience5.8
Honesty-humility6
Extraversion6.7
Agreeableness6.8
Conscientiousness7.3

Agent-ValueBench Schwartz Basic Values (PVQ40)