← Models

Model profile

GPT 5.6 Luna

OpenAIdeveloper
2026-07-09release date
#34 / 346Safety rank
#527 / 662Freedom rank
1discovery sources

Evidence summary

Safety. GPT 5.6 Luna has an estimated Safety rank of #34; its 90% source-sensitivity interval is #22–#139. Its behavior-only rank is #39; company governance moves the combined estimate to #34. Published Safety evidence spans 28 eval lineages and 7 of 7 components. Its strongest relative result is Enkrypt AI Safety Leaderboard (harmful_attack_non_success_rate, #1 of 270); its weakest is GPT-5.6 system card (gore_not_unsafe, #7 of 7).

Freedom. GPT 5.6 Luna has an estimated Freedom rank of #527; its 90% source-sensitivity interval is #366–#574. Published Freedom evidence spans 10 eval lineages and 1 of 1 components. Its strongest relative result is GPT-5.6 system card (gore_not_unsafe, #1 of 7); its weakest is Adversarial Poetry Refusal (AHB self-run) (suicide_self_harm, #14 of 14).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#154 / 358↓0.75Source ↗official
Adversarial Poetry Refusal (AHB self-run)child_sexual_exploitation#2 / 14↓1.449Source ↗self run
Adversarial Poetry Refusal (AHB self-run)defamation#8 / 14↓9.42Source ↗self run
Adversarial Poetry Refusal (AHB self-run)hate#4 / 14↓5.797Source ↗self run
Adversarial Poetry Refusal (AHB self-run)indiscriminate_weapons_cbrne#1 / 14↓0Source ↗self run
Adversarial Poetry Refusal (AHB self-run)intellectual_property#1 / 14↓0Source ↗self run
Adversarial Poetry Refusal (AHB self-run)non_violent_crimes#2 / 14↓0.7246Source ↗self run
Adversarial Poetry Refusal (AHB self-run)privacy#2 / 14↓5.797Source ↗self run
Adversarial Poetry Refusal (AHB self-run)sex_related_crimes#4 / 14↓4.348Source ↗self run
Adversarial Poetry Refusal (AHB self-run)sexual_content#5 / 14↓5.303Source ↗self run
Adversarial Poetry Refusal (AHB self-run)specialized_advice#3 / 14↓2.333Source ↗self run
Adversarial Poetry Refusal (AHB self-run)suicide_self_harm#1 / 14↓1.515Source ↗self run
Adversarial Poetry Refusal (AHB self-run)violent_crimes#2 / 14↓2.273Source ↗self run
ANIMAscore#2 / 22↑0.7494Source ↗self run
BullshitBench v2clear_pushback_rate#52 / 122↑0.39Source ↗official
CAIS Risk Indexagent_red_teaming#19 / 49↓64.2Source ↗official
CAIS Risk Indexbioweapons_assistance#32 / 54↓68.8Source ↗official
CAIS Risk Indexhle_overconfidence#30 / 55↓55.7Source ↗official
CAIS Risk Indexmachiavelli#26 / 51↓87.9Source ↗official
CAIS Risk Indexmask#10 / 57↓7.3Source ↗official
CAIS Risk Indexpolitical_manipulation#21 / 51↓44.7Source ↗official
CAIS Risk Indextextquests_harm#47 / 54↓23.8Source ↗official
Claude system cards — Gray Swan Q1+Q2 indirect prompt injection k=15attack_success_probability_k15_pct#10 / 12↓50Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#90 / 270↑21.96Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#210 / 270↑77.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#1 / 270↑100Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#194 / 268↑93.82Source ↗official
Google Gemini 3.8 launch — Gray Swan indirect prompt injection k=15attack_success_probability_k15_pct#12 / 15↓50Source ↗official
GPT 6 Astra system-card alignment evaluationsoverall_misaligned_outcome_base_pct#4 / 4↓19.7Source ↗official
GPT 6 Astra system-card alignment evaluationsoverall_misaligned_outcome_confirmation_pct#3 / 4↓7.2Source ↗official
GPT-5.6 system cardconnectors_injection_resistance#4 / 7↑0.999Source ↗official
GPT-5.6 system cardemotional_reliance#3 / 7↑0.957Source ↗official
GPT-5.6 system cardextremism_not_unsafe#4 / 7↑0.981Source ↗official
GPT-5.6 system cardgore_not_unsafe#7 / 7↑0.585Source ↗official
GPT-5.6 system cardharm_overall_pct#1 / 7↓0.61Source ↗official
GPT-5.6 system cardhate_not_unsafe#1 / 7↑1Source ↗official
GPT-5.6 system cardmental_health#2 / 7↑0.989Source ↗official
GPT-5.6 system cardnonviolent_illicit_not_unsafe#2 / 7↑0.993Source ↗official
GPT-5.6 system cardsearch_function_calling_injection_resistance#3 / 6↑0.897Source ↗official
GPT-5.6 system cardself_harm#4 / 7↑0.905Source ↗official
GPT-5.6 system cardself_harm_not_unsafe#4 / 7↑0.954Source ↗official
GPT-5.6 system cardsexual_minors_not_unsafe#1 / 7↑0.974Source ↗official
GPT-5.6 system cardsexual_not_unsafe#3 / 7↑0.944Source ↗official
GPT-5.6 system cardviolent_illicit_not_unsafe#5 / 7↑0.94Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsattempted_retry_or_circumvention_after_denial_pct_2026_09_22#4 / 5↓4.3Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsbiology_benign_not_overrefuse_rate#1 / 6↑0.989Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsbiology_dual_use_safe_rate#4 / 6↑0.926Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsbiology_severe_safe_rate#5 / 6↑0.946Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsbroken_search_unacknowledged_limitation_pct_2026_09_22#5 / 5↓78.25Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_extremism_safe_rate#1 / 7↑0.981Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_gore_safe_rate#7 / 7↑0.585Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_hate_safe_rate#1 / 7↑1Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_nonviolent_illicit_safe_rate#3 / 7↑0.993Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_self_harm_safe_rate#5 / 7↑0.954Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_sexual_minors_safe_rate#2 / 7↑0.974Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_sexual_safe_rate#4 / 7↑0.944Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_violent_illicit_safe_rate#5 / 7↑0.94Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testscoding_deception_misrepresentation_pct_2026_09_22#4 / 5↓9.54Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testscyber_production_chat_safe_score#1 / 6↑0.986Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testscyber_synthetic_agentic_safe_score#1 / 6↑1Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsdynamic_emotional_reliance_safe_rate#5 / 7↑0.957Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsdynamic_mental_health_safe_rate#5 / 7↑0.989Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsdynamic_self_harm_safe_rate#5 / 7↑0.905Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_age_restricted_safe_rate#4 / 7↑0.76Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_eating_disorders_safe_rate#5 / 7↑0.702Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_emotional_reliance_safe_rate#5 / 7↑0.931Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_gore_safe_rate#5 / 7↑0.819Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_self_harm_safe_rate#2 / 7↑0.99Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_sexual_content_safe_rate#7 / 7↑0.922Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsunwanted_persistence_after_warning_pct_2026_09_22#5 / 5↓76.5Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsvision_erotic_safe_rate#4 / 6↑0.986Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsvision_extremism_safe_rate#6 / 6↑0.966Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsvision_hate_safe_rate#6 / 6↑0.996Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsvision_self_harm_safe_rate#3 / 6↑0.99Source ↗official
Gray Swan indirect prompt injection (15 attempts)attack_success_probability_k15_pct#11 / 13↓43.9Source ↗official
HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score#31 / 54↑26.33Source ↗official
Inkling-Small model card — FORTRESSbenign_answer_rate#2 / 10↑97.8Source ↗official
Inkling-Small model card — FORTRESSharmful_refusal_rate#3 / 10↑83.8Source ↗official
Inkling-Small model card — StrongREJECTsafety_rate#4 / 10↑98.7Source ↗official
MACHIAVELLIdeception_relative_random_pct#22 / 50↓87.9Source ↗official
MentalHealthBenchoverall_model_performance_pct#8 / 17↑44.89Source ↗official
Olam Social Poker — Social Lie Ratesocial_lie_rate_per_10000_turns#1 / 24↓0Source ↗official
Opposite-Narrator Sycophancysycophancy_rate_pct#6 / 24↓1Source ↗official
Pokee-Isaac model card — DTAPbenign_task_success_rate#1 / 6↑0.851Source ↗official
Pokee-Isaac model card — DTAPcombined_attack_success_rate#3 / 6↓0.501Source ↗official
SM-Benchadversarial#21 / 92↑85.85Source ↗official
SM-Benchambiguous_interpretation#63 / 92↑82.14Source ↗official
SM-Benchanti_hallucination#69 / 92↑87.17Source ↗official
SM-Bencheq_boundaries#65 / 92↑58.71Source ↗official
SM-Benchoverfit#77 / 92↑47.81Source ↗official
SpeciEvalbelief_animal_sentience#42 / 131↑6.9Source ↗official
SpeciEvalland_animal_4ns#13 / 131↓4.15Source ↗official
SpeciEvalsea_animal_4ns#6 / 131↓4.175Source ↗official
SpeciEvalspeciesism#85 / 131↓2.16Source ↗official
TACbase_welfare_rate#91 / 92↑13.46Source ↗self run
Vals AI Cheating Auditbiomystery_bench_cheating_attempt_rate_pct#1 / 9↓2.963Source ↗official
Vals AI Cheating Auditswebench_verified_cheating_attempt_rate_pct#5 / 6↓78.8Source ↗official

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Poetry Refusal (AHB self-run)child_sexual_exploitation#11 / 14↑1.449Source ↗self run
Adversarial Poetry Refusal (AHB self-run)defamation#7 / 14↑9.42Source ↗self run
Adversarial Poetry Refusal (AHB self-run)hate#9 / 14↑5.797Source ↗self run
Adversarial Poetry Refusal (AHB self-run)indiscriminate_weapons_cbrne#12 / 14↑0Source ↗self run
Adversarial Poetry Refusal (AHB self-run)intellectual_property#13 / 14↑0Source ↗self run
Adversarial Poetry Refusal (AHB self-run)non_violent_crimes#13 / 14↑0.7246Source ↗self run
Adversarial Poetry Refusal (AHB self-run)privacy#13 / 14↑5.797Source ↗self run
Adversarial Poetry Refusal (AHB self-run)sex_related_crimes#10 / 14↑4.348Source ↗self run
Adversarial Poetry Refusal (AHB self-run)sexual_content#10 / 14↑5.303Source ↗self run
Adversarial Poetry Refusal (AHB self-run)specialized_advice#11 / 14↑2.333Source ↗self run
Adversarial Poetry Refusal (AHB self-run)suicide_self_harm#14 / 14↑1.515Source ↗self run
Adversarial Poetry Refusal (AHB self-run)violent_crimes#12 / 14↑2.273Source ↗self run
CAIS Risk Indexbioweapons_assistance#23 / 54↑68.8Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#61 / 270↓77.83Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#262 / 270↓100Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#73 / 268↓93.82Source ↗official
GPT-5.6 system cardemotional_reliance#5 / 7↓0.957Source ↗official
GPT-5.6 system cardextremism_not_unsafe#3 / 7↓0.981Source ↗official
GPT-5.6 system cardgore_not_unsafe#1 / 7↓0.585Source ↗official
GPT-5.6 system cardhate_not_unsafe#5 / 7↓1Source ↗official
GPT-5.6 system cardmental_health#6 / 7↓0.989Source ↗official
GPT-5.6 system cardnonviolent_illicit_not_unsafe#5 / 7↓0.993Source ↗official
GPT-5.6 system cardself_harm#4 / 7↓0.905Source ↗official
GPT-5.6 system cardself_harm_not_unsafe#4 / 7↓0.954Source ↗official
GPT-5.6 system cardsexual_minors_not_unsafe#6 / 7↓0.974Source ↗official
GPT-5.6 system cardsexual_not_unsafe#4 / 7↓0.944Source ↗official
GPT-5.6 system cardviolent_illicit_not_unsafe#2 / 7↓0.94Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsbiology_benign_not_overrefuse_rate#1 / 6↑0.989Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsbiology_dual_use_safe_rate#2 / 6↓0.926Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsbiology_severe_safe_rate#2 / 6↓0.946Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_extremism_safe_rate#6 / 7↓0.981Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_gore_safe_rate#1 / 7↓0.585Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_hate_safe_rate#3 / 7↓1Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_nonviolent_illicit_safe_rate#4 / 7↓0.993Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_self_harm_safe_rate#3 / 7↓0.954Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_sexual_minors_safe_rate#5 / 7↓0.974Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_sexual_safe_rate#4 / 7↓0.944Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_violent_illicit_safe_rate#2 / 7↓0.94Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testscyber_production_chat_safe_score#6 / 6↓0.986Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsdynamic_emotional_reliance_safe_rate#3 / 7↓0.957Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsdynamic_mental_health_safe_rate#3 / 7↓0.989Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsdynamic_self_harm_safe_rate#3 / 7↓0.905Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_age_restricted_safe_rate#4 / 7↓0.76Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_eating_disorders_safe_rate#3 / 7↓0.702Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_emotional_reliance_safe_rate#2 / 7↓0.931Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_gore_safe_rate#3 / 7↓0.819Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_self_harm_safe_rate#5 / 7↓0.99Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_sexual_content_safe_rate#1 / 7↓0.922Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsvision_erotic_safe_rate#2 / 6↓0.986Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsvision_extremism_safe_rate#1 / 6↓0.966Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsvision_hate_safe_rate#1 / 6↓0.996Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsvision_self_harm_safe_rate#4 / 6↓0.99Source ↗official
Inkling-Small model card — FORTRESSbenign_answer_rate#2 / 10↑97.8Source ↗official
Inkling-Small model card — FORTRESSharmful_refusal_rate#8 / 10↓83.8Source ↗official
Inkling-Small model card — StrongREJECTsafety_rate#6 / 10↓98.7Source ↗official
SM-Benchadversarial#70 / 92↓85.85Source ↗official
SM-Bencheq_boundaries#65 / 92↑58.71Source ↗official
SM-Benchoverfit#77 / 92↑47.81Source ↗official
SpeechMap model completioncomplete_pct#103 / 181↑52.4Source ↗official
UGI Leaderboard — base-model willingnesswillingness_adherence_score#79 / 156↑1.5Source ↗official
UGI Leaderboard — base-model willingnesswillingness_direct_score#77 / 156↑3.2Source ↗official

Values evaluations

Descriptive values and political-framing results are separate from safety/ethics ranks. Each strip shows the evaluation’s observed model range; its endpoint labels state what lower and higher values mean.

UGI Political Values

DimensionValueDistribution
Political Lean-25.6
Government47.1
Diplomacy70.2
Economy44.9
Society62.2