← Models

Model profile

Glm 5

Z.aideveloper
2026-02-11release date
#47 / 267overall rank
10eval lineages

Evidence summary

Glm 5 has an estimated overall rank of #47; its 90% source-sensitivity interval is #29–#171. Published evidence spans 10 evals and 7 of 7 behavior components. Its strongest relative result is SM-Bench (adversarial, #4 of 73); its weakest is TAC (base_welfare_rate, #60 of 68).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#23 / 3110.34↓ lowerSource ↗official
AgentAbstainabstain#8 / 1761.8↑ higherSource ↗official
AgentAbstaincar#9 / 1759.1↑ higherSource ↗official
AgentAbstainpaired#8 / 1747.8↑ higherSource ↗official
BullshitBench v2clear_pushback_rate#57 / 1050.24↑ higherSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#53 / 10537.79↓ lowerSource ↗official
DystopiaBenchbasaglia_score#9 / 5034.47↓ lowerSource ↗official
DystopiaBenchbaudrillard_score#7 / 5028.97↓ lowerSource ↗official
DystopiaBenchhuxley_score#8 / 5027.5↓ lowerSource ↗official
DystopiaBenchlaguardia_score#8 / 5040.33↓ lowerSource ↗official
DystopiaBenchorwell_score#8 / 5027.87↓ lowerSource ↗official
DystopiaBenchpetrov_score#6 / 5030.8↓ lowerSource ↗official
LiveSecBenchethics#32 / 4331.18↑ higherSource ↗official
LiveSecBenchfactuality#6 / 4375.94↑ higherSource ↗official
LiveSecBenchlegality#25 / 4339.78↑ higherSource ↗official
LiveSecBenchprivacy#18 / 4360.57↑ higherSource ↗official
LiveSecBenchpsychological_health#7 / 4376.17↑ higherSource ↗official
RefusalBenchyouden_j#4 / 190.7291↑ higherSource ↗official
SABERoverall_safety_rate#3 / 1329↑ higherSource ↗official
SABERscenario_a_safety_rate#2 / 1336.25↑ higherSource ↗official
SABERscenario_b_safety_rate#6 / 1333.73↑ higherSource ↗official
SABERscenario_c_safety_rate#6 / 1316.59↑ higherSource ↗official
SM-Benchadversarial#4 / 7390.73↑ higherSource ↗official
SM-Benchambiguous_interpretation#15 / 7389.29↑ higherSource ↗official
SM-Benchanti_hallucination#35 / 7392.67↑ higherSource ↗official
SM-Bencheq_boundaries#40 / 7363.2↑ higherSource ↗official
SM-Benchoverfit#47 / 7363.11↑ higherSource ↗official
TACbase_welfare_rate#60 / 6818.59↑ higherSource ↗self run

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-20.9
Government45.9
Diplomacy65.7
Economy49
Society61.8