← Models

Model profile

Glm 4.7

Z.aideveloper
2026-06-03release date
#138 / 267overall rank
6eval lineages

Evidence summary

Glm 4.7 has an estimated overall rank of #138; its 90% source-sensitivity interval is #59–#254. Its behavior-only rank is #136; company governance moves the combined estimate to #138. Published evidence spans 6 evals and 5 of 7 behavior components. Its strongest relative result is SM-Bench (eq_boundaries, #10 of 73); its weakest is AA-Omniscience (hallucination_rate, #257 of 311).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#257 / 3110.9029↓ lowerSource ↗official
Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate#82 / 10563.96↓ lowerSource ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#21 / 5427.46↑ higherSource ↗official
SABERoverall_safety_rate#9 / 1323.04↑ higherSource ↗official
SABERscenario_a_safety_rate#7 / 1328.03↑ higherSource ↗official
SABERscenario_b_safety_rate#10 / 1326.9↑ higherSource ↗official
SABERscenario_c_safety_rate#10 / 1314.41↑ higherSource ↗official
SM-Benchadversarial#31 / 7382.93↑ higherSource ↗official
SM-Benchambiguous_interpretation#45 / 7382.44↑ higherSource ↗official
SM-Benchanti_hallucination#56 / 7385.86↑ higherSource ↗official
SM-Bencheq_boundaries#10 / 7371.35↑ higherSource ↗official
SM-Benchoverfit#35 / 7373.77↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#33 / 1026.88↑ higherSource ↗official
SpeciEvalland_animal_4ns#37 / 1024.42↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#41 / 1024.7↓ lowerSource ↗official
SpeciEvalspeciesism#77 / 1022.42↓ lowerSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-15.2
Government48
Diplomacy61.2
Economy44.9
Society60.2