Evaluation profile
OpenAI GPT-5.4 User Confirmations
3sub-evals
0.15%total index weight
2components
Within-component eval weight: Human rights 0.462% · Responsible agency 0.541%.
Model score (higher is better)Predicted score
About this eval
OpenAI GPT-5.4 User Confirmations
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| financial_transactionopenai-gpt54/user-confirmations.csv:financial_transactionMeasures financial transaction in the cited first-party or paper evaluation. | human_rights_systemic_harm:0.300|responsible_agency_control:0.700openai-gpt54-system-card | Higher is better | 0.0531% | Human rights 0.163% · Responsible agency 0.191% |
| general_confirmationopenai-gpt54/user-confirmations.csv:general_confirmationMeasures general confirmation in the cited first-party or paper evaluation. | human_rights_systemic_harm:0.300|responsible_agency_control:0.700openai-gpt54-system-card | Higher is better | 0.0443% | Human rights 0.136% · Responsible agency 0.159% |
| high_stakes_communicationopenai-gpt54/user-confirmations.csv:high_stakes_communicationMeasures high stakes communication in the cited first-party or paper evaluation. | human_rights_systemic_harm:0.300|responsible_agency_control:0.700openai-gpt54-system-card | Higher is better | 0.0531% | Human rights 0.163% · Responsible agency 0.191% |
financial_transaction
Measures financial transaction in the cited first-party or paper evaluation.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.2-thinking | 1 | official | |
| 1 | gpt-5.4 | 1 | official | |
| 3 | gpt-5.3-codex | 0.99 | official |
general_confirmation
Measures general confirmation in the cited first-party or paper evaluation.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.2-thinking | 0.94 | official | |
| 1 | gpt-5.4 | 0.94 | official | |
| 3 | gpt-5.3-codex | 0.91 | official |
high_stakes_communication
Measures high stakes communication in the cited first-party or paper evaluation.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-5.2-thinking | 1 | official | |
| 1 | gpt-5.4 | 1 | official | |
| 3 | gpt-5.3-codex | 0.99 | official |