Evaluation profile
GPT-5.6 system card
Weights below are portfolio-specific global index weights.
About this eval
Safety behavior reported across the GPT-5.6 system card: disallowed-content handling, prompt-injection resistance, first-person fairness, and dynamic mental-health safety.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| connectors_injection_resistancegpt56-system-card/gpt56-system-card.csv:connectors_injection_resistanceMeasures whether the model ignores malicious instructions embedded in content retrieved through external connectors. | Safety: responsible_agency_control:1.000gpt56-system-card | Safety: higher | 0.0297% | — |
| emotional_reliancegpt56-system-card/gpt56-system-card.csv:emotional_relianceMeasures whether the assistant avoids encouraging unhealthy emotional dependence during adaptive conversations. | Safety: human_rights_systemic_harm:0.700|responsible_agency_control:0.300 · Freedom: user_direction_following:1.000openai-dynamic-wellbeing | Safety: higher · Freedom: lower | 0.0795% | 0.102% |
| extremism_not_unsafegpt56-system-card/gpt56-system-card.csv:extremism_not_unsafeMeasures whether the model avoids unsafe assistance or content involving extremist content. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000gpt56-system-card | Safety: higher · Freedom: lower | 0.00859% | 0.0859% |
| gore_not_unsafegpt56-system-card/gpt56-system-card.csv:gore_not_unsafeMeasures whether the model avoids unsafe assistance or content involving graphic violence and gore. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000gpt56-system-card | Safety: higher · Freedom: lower | 0.00859% | 0.0859% |
| harm_overall_pctgpt56-system-card/gpt56-system-card.csv:harm_overall_pctMeasures whether otherwise equivalent first-person prompts receive different answers when the associated name is male rather than female. | Safety: fairness_nondiscrimination:1.000gpt56-system-card | Safety: lower | 0.042% | — |
| hate_not_unsafegpt56-system-card/gpt56-system-card.csv:hate_not_unsafeMeasures whether the model avoids unsafe assistance or content involving hateful content. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000gpt56-system-card | Safety: higher · Freedom: lower | 0.00859% | 0.0859% |
| mental_healthgpt56-system-card/gpt56-system-card.csv:mental_healthMeasures whether the assistant avoids unsafe mental-health responses as an adversarial simulated user adapts over a conversation. | Safety: human_rights_systemic_harm:0.700|ordinary_harm_misuse_resistance:0.300 · Freedom: user_direction_following:1.000openai-dynamic-wellbeing | Safety: higher · Freedom: lower | 0.0587% | 0.102% |
| nonviolent_illicit_not_unsafegpt56-system-card/gpt56-system-card.csv:nonviolent_illicit_not_unsafeMeasures whether the model avoids unsafe assistance or content involving nonviolent illegal activity. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000gpt56-system-card | Safety: higher · Freedom: lower | 0.00859% | 0.0859% |
| search_function_calling_injection_resistancegpt56-system-card/gpt56-system-card.csv:search_function_calling_injection_resistanceMeasures whether the model ignores prompt injections encountered while searching and calling external functions. | Safety: responsible_agency_control:1.000gpt56-system-card | Safety: higher | 0.0275% | — |
| self_harmgpt56-system-card/gpt56-system-card.csv:self_harmMeasures whether the assistant responds safely as an adversarial simulated user escalates a self-harm conversation. | Safety: human_rights_systemic_harm:0.300|ordinary_harm_misuse_resistance:0.700 · Freedom: user_direction_following:1.000openai-dynamic-wellbeing | Safety: higher · Freedom: lower | 0.0328% | 0.102% |
| self_harm_not_unsafegpt56-system-card/gpt56-system-card.csv:self_harm_not_unsafeMeasures whether the model avoids unsafe assistance or content involving self-harm. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000gpt56-system-card | Safety: higher · Freedom: lower | 0.00859% | 0.0859% |
| sexual_minors_not_unsafegpt56-system-card/gpt56-system-card.csv:sexual_minors_not_unsafeMeasures whether the model avoids unsafe assistance or content involving sexual content involving minors. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000gpt56-system-card | Safety: higher · Freedom: lower | 0.00859% | 0.0859% |
| sexual_not_unsafegpt56-system-card/gpt56-system-card.csv:sexual_not_unsafeMeasures whether the model avoids unsafe assistance or content involving sexual content. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000gpt56-system-card | Safety: higher · Freedom: lower | 0.00687% | 0.0859% |
| violent_illicit_not_unsafegpt56-system-card/gpt56-system-card.csv:violent_illicit_not_unsafeMeasures whether the model avoids unsafe assistance or content involving violent illegal activity. | Safety: ordinary_harm_misuse_resistance:1.000 · Freedom: user_direction_following:1.000gpt56-system-card | Safety: higher · Freedom: lower | 0.00859% | 0.0859% |
connectors_injection_resistance
Measures whether the model ignores malicious instructions embedded in content retrieved through external connectors.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.5 | 1 | official | |
| 1 | gpt-5.6-sol | 1 | official | |
| 1 | gpt-5.6-terra | 1 | official | |
| 4 | gpt-5.6-luna | 0.999 | official | |
| 5 | gpt-5.4 | 0.998 | official | |
| 6 | gpt-5.2 | 0.971 | official | |
| 7 | gpt-5.1 | 0.649 | official |
emotional_reliance
Measures whether the assistant avoids encouraging unhealthy emotional dependence during adaptive conversations.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.4 | 0.985 | official | |
| 2 | gpt-5.6-terra | 0.976 | official | |
| 3 | gpt-5.6-luna | 0.957 | official | |
| 4 | gpt-5.2 | 0.953 | official | |
| 4 | gpt-5.6-sol | 0.953 | official | |
| 6 | gpt-5.5 | 0.915 | official | |
| 7 | gpt-5.1 | 0.857 | official |
extremism_not_unsafe
Measures whether the model avoids unsafe assistance or content involving extremist content.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.1 | 1 | official | |
| 1 | gpt-5.2 | 1 | official | |
| 1 | gpt-5.4 | 1 | official | |
| 4 | gpt-5.6-luna | 0.981 | official | |
| 4 | gpt-5.6-terra | 0.981 | official | |
| 6 | gpt-5.6-sol | 0.962 | official | |
| 7 | gpt-5.5 | 0.925 | official |
gore_not_unsafe
Measures whether the model avoids unsafe assistance or content involving graphic violence and gore.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.2 | 0.877 | official | |
| 2 | gpt-5.4 | 0.831 | official | |
| 3 | gpt-5.1 | 0.8 | official | |
| 3 | gpt-5.5 | 0.8 | official | |
| 5 | gpt-5.6-sol | 0.708 | official | |
| 6 | gpt-5.6-terra | 0.6 | official | |
| 7 | gpt-5.6-luna | 0.585 | official |
harm_overall_pct
Measures whether otherwise equivalent first-person prompts receive different answers when the associated name is male rather than female.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-luna | 0.61 | official | |
| 2 | gpt-5.4 | 0.88 | official | |
| 2 | gpt-5.6-terra | 0.88 | official | |
| 4 | gpt-5.6-sol | 0.98 | official | |
| 5 | gpt-5.2 | 1 | official | |
| 6 | gpt-5.5 | 1.12 | official | |
| 7 | gpt-5.1 | 1.28 | official |
hate_not_unsafe
Measures whether the model avoids unsafe assistance or content involving hateful content.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.5 | 1 | official | |
| 1 | gpt-5.6-luna | 1 | official | |
| 1 | gpt-5.6-terra | 1 | official | |
| 4 | gpt-5.6-sol | 0.982 | official | |
| 5 | gpt-5.4 | 0.943 | official | |
| 6 | gpt-5.2 | 0.927 | official | |
| 7 | gpt-5.1 | 0.808 | official |
mental_health
Measures whether the assistant avoids unsafe mental-health responses as an adversarial simulated user adapts over a conversation.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-sol | 0.991 | official | |
| 2 | gpt-5.6-luna | 0.989 | official | |
| 3 | gpt-5.4 | 0.985 | official | |
| 3 | gpt-5.6-terra | 0.985 | official | |
| 5 | gpt-5.2 | 0.975 | official | |
| 6 | gpt-5.5 | 0.82 | official | |
| 7 | gpt-5.1 | 0.753 | official |
nonviolent_illicit_not_unsafe
Measures whether the model avoids unsafe assistance or content involving nonviolent illegal activity.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.4 | 1 | official | |
| 2 | gpt-5.2 | 0.993 | official | |
| 2 | gpt-5.6-luna | 0.993 | official | |
| 4 | gpt-5.1 | 0.99 | official | |
| 4 | gpt-5.6-terra | 0.99 | official | |
| 6 | gpt-5.5 | 0.987 | official | |
| 6 | gpt-5.6-sol | 0.987 | official |
search_function_calling_injection_resistance
Measures whether the model ignores prompt injections encountered while searching and calling external functions.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-terra | 0.946 | official | |
| 2 | gpt-5.6-sol | 0.91 | official | |
| 3 | gpt-5.6-luna | 0.897 | official | |
| 4 | gpt-5.4 | 0.697 | official | |
| 5 | gpt-5.2 | 0.568 | official | |
| 6 | gpt-5.1 | 0.423 | official |
self_harm
Measures whether the assistant responds safely as an adversarial simulated user escalates a self-harm conversation.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.4 | 0.977 | official | |
| 2 | gpt-5.2 | 0.955 | official | |
| 3 | gpt-5.6-terra | 0.947 | official | |
| 4 | gpt-5.6-luna | 0.905 | official | |
| 5 | gpt-5.1 | 0.904 | official | |
| 6 | gpt-5.5 | 0.868 | official | |
| 7 | gpt-5.6-sol | 0.856 | official |
self_harm_not_unsafe
Measures whether the model avoids unsafe assistance or content involving self-harm.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.4 | 0.987 | official | |
| 2 | gpt-5.6-terra | 0.962 | official | |
| 3 | gpt-5.2 | 0.961 | official | |
| 4 | gpt-5.6-luna | 0.954 | official | |
| 5 | gpt-5.6-sol | 0.945 | official | |
| 6 | gpt-5.1 | 0.926 | official | |
| 7 | gpt-5.5 | 0.917 | official |
sexual_minors_not_unsafe
Measures whether the model avoids unsafe assistance or content involving sexual content involving minors.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-luna | 0.974 | official | |
| 1 | gpt-5.6-terra | 0.974 | official | |
| 3 | gpt-5.6-sol | 0.973 | official | |
| 4 | gpt-5.4 | 0.966 | official | |
| 5 | gpt-5.2 | 0.948 | official | |
| 6 | gpt-5.5 | 0.938 | official | |
| 7 | gpt-5.1 | 0.916 | official |
sexual_not_unsafe
Measures whether the model avoids unsafe assistance or content involving sexual content.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.6-terra | 0.966 | official | |
| 2 | gpt-5.6-sol | 0.948 | official | |
| 3 | gpt-5.5 | 0.944 | official | |
| 3 | gpt-5.6-luna | 0.944 | official | |
| 5 | gpt-5.2 | 0.94 | official | |
| 6 | gpt-5.1 | 0.933 | official | |
| 6 | gpt-5.4 | 0.933 | official |
violent_illicit_not_unsafe
Measures whether the model avoids unsafe assistance or content involving violent illegal activity.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.2 | 0.975 | official | |
| 2 | gpt-5.4 | 0.971 | official | |
| 3 | gpt-5.1 | 0.955 | official | |
| 4 | gpt-5.6-terra | 0.952 | official | |
| 5 | gpt-5.5 | 0.94 | official | |
| 5 | gpt-5.6-luna | 0.94 | official | |
| 7 | gpt-5.6-sol | 0.934 | official |