← Models

Model profile

Granite 4.0 Micro

IBMdeveloper
2025-10-02release date
#234 / 267overall rank
9eval lineages

Evidence summary

Granite 4.0 Micro has an estimated overall rank of #234; its 90% source-sensitivity interval is #72–#237. Its behavior-only rank is #238; company governance moves the combined estimate to #234. Published evidence spans 9 evals and 6 of 7 behavior components. Its strongest relative result is Enkrypt AI Safety Leaderboard (toxicity_attack_non_success_rate, #31 of 258); its weakest is UAVBench safety-critical decision recognition (ethical_safety_critical_accuracy, #25 of 27).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#115 / 3110.7447↓ lowerSource ↗official
AIRBench 2024 Safety Scenariossafety_scenarios#42 / 800.661↑ higherSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#141 / 26014.99↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#94 / 26090.17↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#59 / 26088.89↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#31 / 25899.45↑ higherSource ↗official
HELM Safetyanthropic_red_team#44 / 800.987↑ higherSource ↗official
HELM Safetybbq#69 / 800.779↑ higherSource ↗official
HELM Safetyharmbench#47 / 800.695↑ higherSource ↗official
HELM Safetysimple_safety_tests#68 / 800.945↑ higherSource ↗official
HELM Safetyxstest#73 / 800.869↑ higherSource ↗official
UAVBench safety-critical decision recognitionethical_safety_critical_accuracy#25 / 270.54↑ higherSource ↗official