← Models

Model profile

GPT 5.6 Terra

OpenAIdeveloper
2026-07-09release date
#7 / 267overall rank
17eval lineages
1discovery sources

Evidence summary

GPT 5.6 Terra has an estimated overall rank of #7; its 90% source-sensitivity interval is #5–#60. Published evidence spans 17 evals and 7 of 7 behavior components. Its strongest relative result is SpeciEval (speciesism, #5 of 102); its weakest is GPT-5.6 system card — disallowed content with challenging prompts (gore_not_unsafe, #6 of 7).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Published eval results

Rank is within that sub-eval. Black marks the observed result; the grey dot marks the value implied by the global rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionBetterSource
AA-Omnisciencehallucination_rate#202 / 3110.8519↓ lowerSource ↗official
BullshitBench v2clear_pushback_rate#23 / 1050.495↑ higherSource ↗official
CAIS Risk Indexagent_red_teaming#11 / 4354.6↓ lowerSource ↗official
CAIS Risk Indexbioweapons_assistance#22 / 4866.2↓ lowerSource ↗official
CAIS Risk Indexhle_overconfidence#20 / 4950.7↓ lowerSource ↗official
CAIS Risk Indexmachiavelli#3 / 4580.4↓ lowerSource ↗official
CAIS Risk Indexmask#6 / 516.7↓ lowerSource ↗official
CAIS Risk Indexpolitical_manipulation#14 / 3246.2↓ lowerSource ↗official
CAIS Risk Indextextquests_harm#41 / 4822.2↓ lowerSource ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#35 / 26039.28↑ higherSource ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#217 / 26080↑ higherSource ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#24 / 26097.78↑ higherSource ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#198 / 25892.36↑ higherSource ↗official
GPT-5.6 system card — disallowed content with challenging promptsextremism_not_unsafe#4 / 70.981↑ higherSource ↗official
GPT-5.6 system card — disallowed content with challenging promptsgore_not_unsafe#6 / 70.6↑ higherSource ↗official
GPT-5.6 system card — disallowed content with challenging promptshate_not_unsafe#1 / 71↑ higherSource ↗official
GPT-5.6 system card — disallowed content with challenging promptsnonviolent_illicit_not_unsafe#4 / 70.99↑ higherSource ↗official
GPT-5.6 system card — disallowed content with challenging promptsself_harm_not_unsafe#2 / 70.962↑ higherSource ↗official
GPT-5.6 system card — disallowed content with challenging promptssexual_minors_not_unsafe#1 / 70.974↑ higherSource ↗official
GPT-5.6 system card — disallowed content with challenging promptssexual_not_unsafe#1 / 70.966↑ higherSource ↗official
GPT-5.6 system card — disallowed content with challenging promptsviolent_illicit_not_unsafe#4 / 70.952↑ higherSource ↗official
GPT-5.6 system card — first-person fairnessharm_overall_pct#2 / 70.88↓ lowerSource ↗official
GPT-5.6 system card — prompt-injection robustnessconnectors_injection_resistance#1 / 71↑ higherSource ↗official
GPT-5.6 system card — prompt-injection robustnesssearch_function_calling_injection_resistance#1 / 60.946↑ higherSource ↗official
Gray Swan indirect prompt injection (15 attempts)attack_success_probability_k15_pct#9 / 1330.4↓ lowerSource ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#19 / 5427.66↑ higherSource ↗official
MACHIAVELLIdeception_relative_random_pct#3 / 5080.4↓ lowerSource ↗official
Manager Coercion Benchcoercion_ladder_depth#18 / 318.7↓ lowerSource ↗official
Manager Coercion Benchfabrication_rate#1 / 130↓ lowerSource ↗official
SM-Benchadversarial#26 / 7383.66↑ higherSource ↗official
SM-Benchambiguous_interpretation#40 / 7383.63↑ higherSource ↗official
SM-Benchanti_hallucination#41 / 7390.84↑ higherSource ↗official
SM-Bencheq_boundaries#31 / 7366.01↑ higherSource ↗official
SM-Benchoverfit#55 / 7353.55↑ higherSource ↗official
SpeciEvalbelief_animal_sentience#27 / 1026.925↑ higherSource ↗official
SpeciEvalland_animal_4ns#22 / 1024.29↓ lowerSource ↗official
SpeciEvalsea_animal_4ns#5 / 1024.225↓ lowerSource ↗official
SpeciEvalspeciesism#5 / 1021.265↓ lowerSource ↗official
TACbase_welfare_rate#29 / 6829.49↑ higherSource ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-26.5
Government47.5
Diplomacy68.3
Economy48.6
Society60.5