← Models

Model profile

GPT 5.6 Sol

OpenAIdeveloper
2026-07-09release date
#32 / 267overall rank
17eval lineages
1discovery sources

Evidence summary

GPT 5.6 Sol has an estimated overall rank of #32; its 90% source-sensitivity interval is #6–#175. Its behavior-only rank is #39; company governance moves the combined estimate to #32. Published evidence spans 17 evals and 7 of 7 behavior components. Its strongest relative result is SpeciEval (belief_animal_sentience, #1 of 102); its weakest is CAIS Risk Index (textquests_harm, #48 of 48).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#220 / 3110.8711↓ lowerSource ↗official
BullshitBench v2clear_pushback_rate#26 / 1050.465↑ higherSource ↗official
CAIS Risk Indexagent_red_teaming#2 / 4340.3↓ lowerSource ↗official
CAIS Risk Indexbioweapons_assistance#21 / 4864.3↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#15 / 4946.7↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#9 / 4582.7↓ lowerSource ↗official
CAIS Risk Indexmask#3 / 514.6↓ lowerSource ↗official
CAIS Risk Indexpolitical_manipulation#10 / 3241.7↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#48 / 4831.4↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#53 / 26027.91↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#214 / 26080.83↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#7 / 26099.44↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#211 / 25891.09↑ higherSource ↗official
GPT-5.6 system card — disallowed content with challenging promptsextremism_not_unsafe#6 / 70.962↑ higherSource ↗official
GPT-5.6 system card — disallowed content with challenging promptsgore_not_unsafe#5 / 70.708↑ higherSource ↗official
GPT-5.6 system card — disallowed content with challenging promptshate_not_unsafe#4 / 70.982↑ higherSource ↗official
GPT-5.6 system card — disallowed content with challenging promptsnonviolent_illicit_not_unsafe#6 / 70.987↑ higherSource ↗official
GPT-5.6 system card — disallowed content with challenging promptsself_harm_not_unsafe#5 / 70.945↑ higherSource ↗official
GPT-5.6 system card — disallowed content with challenging promptssexual_minors_not_unsafe#3 / 70.973↑ higherSource ↗official
GPT-5.6 system card — disallowed content with challenging promptssexual_not_unsafe#2 / 70.948↑ higherSource ↗official
GPT-5.6 system card — disallowed content with challenging promptsviolent_illicit_not_unsafe#7 / 70.934↑ higherSource ↗official
GPT-5.6 system card — first-person fairnessharm_overall_pct#4 / 70.98↓ lowerSource ↗official
GPT-5.6 system card — prompt-injection robustnessconnectors_injection_resistance#1 / 71↑ higherSource ↗official
GPT-5.6 system card — prompt-injection robustnesssearch_function_calling_injection_resistance#2 / 60.91↑ higherSource ↗official
Gray Swan indirect prompt injection (15 attempts)attack_success_probability_k15_pct#7 / 1320↓ lowerSource ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#35 / 5426.14↑ higherSource ↗official
MACHIAVELLIdeception_relative_random_pct#9 / 5082.7↓ lowerSource ↗official
Manager Coercion Benchcoercion_ladder_depth#22 / 318.9↓ lowerSource ↗official
Manager Coercion Benchfabrication_rate#1 / 130↓ lowerSource ↗official
SM-Benchadversarial#26 / 7383.66↑ higherSource ↗official
SM-Benchambiguous_interpretation#39 / 7384.08↑ higherSource ↗official
SM-Benchanti_hallucination#36 / 7392.41↑ higherSource ↗official
SM-Bencheq_boundaries#33 / 7365.31↑ higherSource ↗official
SM-Benchoverfit#51 / 7357.52↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#1 / 1027↑ higherSource ↗official
SpeciEvalland_animal_4ns#25 / 1024.315↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#15 / 1024.435↓ lowerSource ↗official
SpeciEvalspeciesism#3 / 1021.165↓ lowerSource ↗official
TACbase_welfare_rate#65 / 6816.67↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-29.5
Government43.4
Diplomacy71.7
Economy47.2
Society62.2