← Evals

Evaluation profile

OpenAI GPT-5.3 Dynamic Wellbeing

3sub-evals
0.464%total index weight
3components

Within-component eval weight: Human rights 2.6% · Responsible agency 0.256% · Misuse resistance 0.346%.

Model score (higher is better)Predicted score

About this eval

OpenAI GPT-5.3 Dynamic Wellbeing

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
emotional_relianceopenai-gpt53/dynamic-wellbeing.csv:emotional_relianceMeasures emotional reliance in the cited first-party or paper evaluation.human_rights_systemic_harm:0.700|responsible_agency_control:0.300openai-gpt53-system-cardHigher is better0.217%Human rights 1.19% · Responsible agency 0.256%
mental_healthopenai-gpt53/dynamic-wellbeing.csv:mental_healthMeasures mental health in the cited first-party or paper evaluation.human_rights_systemic_harm:0.700|ordinary_harm_misuse_resistance:0.300openai-gpt53-system-cardHigher is better0.159%Human rights 0.991% · Misuse resistance 0.104%
self_harmopenai-gpt53/dynamic-wellbeing.csv:self_harmMeasures self harm in the cited first-party or paper evaluation.human_rights_systemic_harm:0.300|ordinary_harm_misuse_resistance:0.700openai-gpt53-system-cardHigher is better0.0879%Human rights 0.425% · Misuse resistance 0.242%

emotional_reliance

Measures emotional reliance in the cited first-party or paper evaluation.

RankModelValueRelative performanceProvenance
1gpt-5.3-instant0.992official
2gpt-5.2-instant0.952official
3gpt-5.10.945official

mental_health

Measures mental health in the cited first-party or paper evaluation.

RankModelValueRelative performanceProvenance
1gpt-5.2-instant1official
2gpt-5.3-instant0.985official
3gpt-5.10.832official

self_harm

Measures self harm in the cited first-party or paper evaluation.

RankModelValueRelative performanceProvenance
1gpt-5.2-instant0.92official
2gpt-5.3-instant0.911official
3gpt-5.10.845official