← Models

Model profile

Glm 4.5

Z.aideveloper
2025-07-28release date
#170 / 267overall rank
9eval lineages

Evidence summary

Glm 4.5 has an estimated overall rank of #170; its 90% source-sensitivity interval is #68–#219. Its behavior-only rank is #164; company governance moves the combined estimate to #170. Published evidence spans 9 evals and 7 of 7 behavior components. Its strongest relative result is Confabulations (confabulation_rate, #8 of 52); its weakest is FlagEval Safety and Values (a1_qualified_rate, #18 of 18).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#91 / 3110.6716↓ lowerSource ↗official
BullshitBench v2clear_pushback_rate#93 / 1050.07↑ higherSource ↗official
Confabulationsconfabulation_rate#8 / 527.921↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#49 / 26030.49↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#228 / 26077.67↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#190 / 26052.22↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#145 / 25895.76↑ higherSource ↗official
FlagEval Safety and Valuesa1_qualified_rate#18 / 1859.1↑ higherSource ↗official
FlagEval Safety and Valuesa2_qualified_rate#18 / 1857.84↑ higherSource ↗official
FlagEval Safety and Valuesa3_qualified_rate#18 / 1860.32↑ higherSource ↗official
FlagEval Safety and Valuesa4_qualified_rate#18 / 1860.7↑ higherSource ↗official
FlagEval Safety and Valuesa5_qualified_rate#17 / 1856.79↑ higherSource ↗official
FORTRESSaverage_risk_score#44 / 4959.58↓ lowerSource ↗official
FORTRESSover_refusal_score#15 / 462.54↓ lowerSource ↗official
MASKlying_probability_pct#25 / 5338.54↓ lowerSource ↗official
Social Welfare Function Benchmarkfairness#13 / 190.475↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#61 / 1026.7↑ higherSource ↗official
SpeciEvalland_animal_4ns#42 / 1024.47↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#88 / 1025.09↓ lowerSource ↗official
SpeciEvalspeciesism#20 / 1021.58↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-18.1
Government46.5
Diplomacy66.2
Economy45.2
Society60.8