← Models

Model profile

Glm 5.2

Z.aideveloper
2026-06-16release date
#86 / 267overall rank
15eval lineages
4discovery sources

Evidence summary

Glm 5.2 has an estimated overall rank of #86; its 90% source-sensitivity interval is #45–#188. Its behavior-only rank is #78; company governance moves the combined estimate to #86. Published evidence spans 15 evals and 6 of 7 behavior components. Its strongest relative result is SpeciEval (belief_animal_sentience, #1 of 102); its weakest is TAC (base_welfare_rate, #57 of 68).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#12 / 3110.2813↓ lowerSource ↗official
ANIMAscore#8 / 190.7101↑ higherSource ↗official
BullshitBench v2clear_pushback_rate#57 / 1050.24↑ higherSource ↗official
CAIS Risk Indexagent_red_teaming#19 / 4371↓ lowerSource ↗official
CAIS Risk Indexbioweapons_assistance#17 / 4863.8↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#8 / 4942↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#31 / 4592.3↓ lowerSource ↗official
CAIS Risk Indexmask#30 / 5125.8↓ lowerSource ↗official
CAIS Risk Indexpolitical_manipulation#19 / 3250↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#25 / 4818.1↓ lowerSource ↗official
MACHIAVELLIdeception_relative_random_pct#32 / 5092.3↓ lowerSource ↗official
Manager Coercion Benchcoercion_ladder_depth#21 / 318.867↓ lowerSource ↗self run
PHAREbias_resistance_diagnostic#23 / 660.5127↑ higherSource ↗official
PHAREhallucination_resistance_diagnostic#25 / 700.7823↑ higherSource ↗official
PHAREharm_resistance_diagnostic#34 / 700.9373↑ higherSource ↗official
PHAREjailbreak_resistance_diagnostic#27 / 670.5646↑ higherSource ↗official
SM-Benchadversarial#12 / 7386.83↑ higherSource ↗official
SM-Benchambiguous_interpretation#27 / 7387.2↑ higherSource ↗official
SM-Benchanti_hallucination#22 / 7396.34↑ higherSource ↗official
SM-Bencheq_boundaries#38 / 7364.04↑ higherSource ↗official
SM-Benchoverfit#25 / 7380.33↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#1 / 1027↑ higherSource ↗official
SpeciEvalland_animal_4ns#72 / 1024.75↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#58 / 1024.8↓ lowerSource ↗official
SpeciEvalspeciesism#80 / 1022.48↓ lowerSource ↗official
TACbase_welfare_rate#57 / 6819.23↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-14.2
Government45.8
Diplomacy62.1
Economy47.4
Society60.4