← Models

Model profile

GPT-6 Luna

OpenAIdeveloper
2026-09-22release date
#8 / 346Safety rank
#612 / 662Freedom rank

Evidence summary

Safety. GPT-6 Luna has an estimated Safety rank of #8; its 90% source-sensitivity interval is #3–#72. Its behavior-only rank is #10; company governance moves the combined estimate to #8. Published Safety evidence spans 12 eval lineages and 7 of 7 components. Its strongest relative result is Enkrypt AI Safety Leaderboard (harmful_attack_non_success_rate, #1 of 270); its weakest is GPT-6 Sol/Luna system card — safety and updated alignment tests (agentic_chat_plugins_safe_rate, #4 of 4).

Freedom. GPT-6 Luna has an estimated Freedom rank of #612; its 90% source-sensitivity interval is #371–#644. Published Freedom evidence spans 5 eval lineages and 1 of 1 components. Its strongest relative result is Enkrypt AI Safety Leaderboard (cbrn_attack_non_success_rate, #32 of 270); its weakest is Adversarial Poetry Refusal (AHB self-run) (child_sexual_exploitation, #14 of 14).

Compare this model

Only models sharing at least one published sub-eval are listed.

Official and reference links

Safety evals

Rank and direction are specific to the Safety portfolio. Black marks the observed result; the grey dot marks the value implied by the Safety rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
AA-Omnisciencehallucination_rate#161 / 358↓0.7673Source ↗official
Adversarial Poetry Refusal (AHB self-run)child_sexual_exploitation#1 / 14↓0Source ↗self run
Adversarial Poetry Refusal (AHB self-run)defamation#1 / 14↓0Source ↗self run
Adversarial Poetry Refusal (AHB self-run)hate#1 / 14↓1.449Source ↗self run
Adversarial Poetry Refusal (AHB self-run)indiscriminate_weapons_cbrne#4 / 14↓0.8333Source ↗self run
Adversarial Poetry Refusal (AHB self-run)intellectual_property#1 / 14↓0Source ↗self run
Adversarial Poetry Refusal (AHB self-run)non_violent_crimes#1 / 14↓0Source ↗self run
Adversarial Poetry Refusal (AHB self-run)privacy#1 / 14↓0.7246Source ↗self run
Adversarial Poetry Refusal (AHB self-run)sex_related_crimes#1 / 14↓1.481Source ↗self run
Adversarial Poetry Refusal (AHB self-run)sexual_content#2 / 14↓3.03Source ↗self run
Adversarial Poetry Refusal (AHB self-run)specialized_advice#1 / 14↓0.6734Source ↗self run
Adversarial Poetry Refusal (AHB self-run)suicide_self_harm#2 / 14↓1.587Source ↗self run
Adversarial Poetry Refusal (AHB self-run)violent_crimes#1 / 14↓0.7576Source ↗self run
BullshitBench v2clear_pushback_rate#26 / 122↑0.5825Source ↗official
Enkrypt AI Safety Leaderboardbias_attack_non_success_rate#44 / 270↑37.98Source ↗official
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#238 / 270↑67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#1 / 270↑100Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#205 / 268↑92.91Source ↗official
FORTRESSaverage_risk_score#2 / 60↓10.78Source ↗official
FORTRESSover_refusal_score#43 / 59↓7.62Source ↗official
GPT 6 Astra system-card alignment evaluationsfollowed_unauthorized_agent_instructions_pct#1 / 4↓0Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsagentic_chat_plugins_safe_rate#4 / 4↑0.846Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsagentic_codex_age_restricted_safe_rate#3 / 4↑0.687Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsagentic_codex_nonviolent_wrongdoing_safe_rate#1 / 4↑1Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsagentic_codex_personal_data_safe_rate#2 / 4↑0.795Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsagentic_codex_self_harm_safe_rate#1 / 4↑0.96Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsagentic_codex_violent_wrongdoing_safe_rate#3 / 4↑0.796Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsagentic_redteam_chat_plugins_safe_rate#3 / 4↑0.814Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsagentic_redteam_codex_safe_rate#2 / 4↑0.957Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsattempted_retry_or_circumvention_after_denial_pct_2026_09_22#3 / 5↓0.3Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsbiology_benign_not_overrefuse_rate#5 / 6↑0.958Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsbiology_dual_use_safe_rate#3 / 6↑0.962Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsbiology_severe_safe_rate#3 / 6↑0.985Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsbroken_search_unacknowledged_limitation_pct_2026_09_22#3 / 5↓28.67Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_extremism_safe_rate#3 / 7↑0.979Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_gore_safe_rate#3 / 7↑0.876Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_hate_safe_rate#1 / 7↑1Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_nonviolent_illicit_safe_rate#1 / 7↑0.997Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_self_harm_safe_rate#3 / 7↑0.984Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_sexual_minors_safe_rate#5 / 7↑0.965Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_sexual_safe_rate#3 / 7↑0.962Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_violent_illicit_safe_rate#3 / 7↑0.986Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testscoding_deception_misrepresentation_pct_2026_09_22#3 / 5↓2.81Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testscyber_production_chat_safe_score#5 / 6↑0.951Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testscyber_semisynthetic_agentic_safe_score#2 / 5↑0.987Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testscyber_synthetic_agentic_safe_score#4 / 6↑0.997Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsdynamic_emotional_reliance_safe_rate#4 / 7↑0.963Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsdynamic_mental_health_safe_rate#1 / 7↑1Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsdynamic_self_harm_safe_rate#4 / 7↑0.924Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsinstruction_hierarchy_defender_success_pct#2 / 3↑99.97Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsstatic_jailbreak_bio_high_risk_defender_success_pct#3 / 5↑73.8Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsstatic_jailbreak_bio_severe_defender_success_pct#3 / 5↑73Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsstatic_jailbreak_cyber_defender_success_pct#2 / 5↑87.5Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsstatic_jailbreak_violence_moderate_defender_success_pct#3 / 5↑77Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsstatic_jailbreak_violence_severe_defender_success_pct#3 / 5↑89.3Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_age_restricted_safe_rate#3 / 7↑0.848Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_eating_disorders_safe_rate#2 / 7↑0.871Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_emotional_reliance_safe_rate#2 / 7↑0.946Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_gore_safe_rate#3 / 7↑0.878Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_self_harm_safe_rate#5 / 7↑0.982Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_sexual_content_safe_rate#3 / 7↑0.949Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsunwanted_persistence_after_warning_pct_2026_09_22#2 / 5↓42.4Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsvision_erotic_safe_rate#2 / 6↑0.995Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsvision_extremism_safe_rate#3 / 6↑0.982Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsvision_hate_safe_rate#3 / 6↑0.998Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsvision_self_harm_safe_rate#1 / 6↑0.999Source ↗official
Manager Coercion Benchcoercion_ladder_depth#6 / 45↓4.433Source ↗self run
MentalHealthBenchoverall_model_performance_pct#4 / 17↑50.21Source ↗official
Olam Diplomacy — Broken Promise Ratebroken_promise_rate_pct#2 / 19↓10.5Source ↗official
SM-Benchadversarial#75 / 92↑77.56Source ↗official
SM-Benchambiguous_interpretation#77 / 92↑77.08Source ↗official
SM-Benchanti_hallucination#52 / 92↑92.67Source ↗official
SM-Bencheq_boundaries#49 / 92↑64.33Source ↗official
SM-Benchoverfit#75 / 92↑49.18Source ↗official
SpeciEvalbelief_animal_sentience#36 / 131↑6.93Source ↗official
SpeciEvalland_animal_4ns#11 / 131↓4.08Source ↗official
SpeciEvalsea_animal_4ns#3 / 131↓3.92Source ↗official
SpeciEvalspeciesism#51 / 131↓1.85Source ↗official
TACbase_welfare_rate#48 / 92↑26.28Source ↗self run

Freedom evals

Rank and direction are specific to the Freedom portfolio. Black marks the observed result; the grey dot marks the value implied by the Freedom rank. Values stay on each source’s native scale.

Eval / sub-evalRankValueDistributionSource
Adversarial Poetry Refusal (AHB self-run)child_sexual_exploitation#14 / 14↑0Source ↗self run
Adversarial Poetry Refusal (AHB self-run)defamation#14 / 14↑0Source ↗self run
Adversarial Poetry Refusal (AHB self-run)hate#14 / 14↑1.449Source ↗self run
Adversarial Poetry Refusal (AHB self-run)indiscriminate_weapons_cbrne#11 / 14↑0.8333Source ↗self run
Adversarial Poetry Refusal (AHB self-run)intellectual_property#13 / 14↑0Source ↗self run
Adversarial Poetry Refusal (AHB self-run)non_violent_crimes#14 / 14↑0Source ↗self run
Adversarial Poetry Refusal (AHB self-run)privacy#14 / 14↑0.7246Source ↗self run
Adversarial Poetry Refusal (AHB self-run)sex_related_crimes#14 / 14↑1.481Source ↗self run
Adversarial Poetry Refusal (AHB self-run)sexual_content#13 / 14↑3.03Source ↗self run
Adversarial Poetry Refusal (AHB self-run)specialized_advice#14 / 14↑0.6734Source ↗self run
Adversarial Poetry Refusal (AHB self-run)suicide_self_harm#12 / 14↑1.587Source ↗self run
Adversarial Poetry Refusal (AHB self-run)violent_crimes#14 / 14↑0.7576Source ↗self run
Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate#32 / 270↓67Source ↗official
Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate#262 / 270↓100Source ↗official
Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate#63 / 268↓92.91Source ↗official
FORTRESSaverage_risk_score#59 / 60↑10.78Source ↗official
FORTRESSover_refusal_score#43 / 59↓7.62Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsbiology_benign_not_overrefuse_rate#5 / 6↑0.958Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsbiology_dual_use_safe_rate#4 / 6↓0.962Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsbiology_severe_safe_rate#4 / 6↓0.985Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_extremism_safe_rate#5 / 7↓0.979Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_gore_safe_rate#5 / 7↓0.876Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_hate_safe_rate#3 / 7↓1Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_nonviolent_illicit_safe_rate#6 / 7↓0.997Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_self_harm_safe_rate#5 / 7↓0.984Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_sexual_minors_safe_rate#3 / 7↓0.965Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_sexual_safe_rate#5 / 7↓0.962Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_violent_illicit_safe_rate#5 / 7↓0.986Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testscyber_production_chat_safe_score#2 / 6↓0.951Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsdynamic_emotional_reliance_safe_rate#4 / 7↓0.963Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsdynamic_mental_health_safe_rate#6 / 7↓1Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsdynamic_self_harm_safe_rate#4 / 7↓0.924Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsinstruction_hierarchy_defender_success_pct#1 / 3↓99.97Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsstatic_jailbreak_bio_high_risk_defender_success_pct#3 / 5↓73.8Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsstatic_jailbreak_bio_severe_defender_success_pct#3 / 5↓73Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsstatic_jailbreak_cyber_defender_success_pct#4 / 5↓87.5Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsstatic_jailbreak_violence_moderate_defender_success_pct#3 / 5↓77Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsstatic_jailbreak_violence_severe_defender_success_pct#3 / 5↓89.3Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_age_restricted_safe_rate#5 / 7↓0.848Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_eating_disorders_safe_rate#6 / 7↓0.871Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_emotional_reliance_safe_rate#6 / 7↓0.946Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_gore_safe_rate#5 / 7↓0.878Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_self_harm_safe_rate#2 / 7↓0.982Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsu18_sexual_content_safe_rate#5 / 7↓0.949Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsvision_erotic_safe_rate#5 / 6↓0.995Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsvision_extremism_safe_rate#4 / 6↓0.982Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsvision_hate_safe_rate#2 / 6↓0.998Source ↗official
GPT-6 Sol/Luna system card — safety and updated alignment testsvision_self_harm_safe_rate#6 / 6↓0.999Source ↗official
SM-Benchadversarial#16 / 92↓77.56Source ↗official
SM-Bencheq_boundaries#49 / 92↑64.33Source ↗official
SM-Benchoverfit#75 / 92↑49.18Source ↗official