← Evals

Evaluation profile

GPT-5.6 system card — prompt-injection robustness

2sub-evals
0.0806%total index weight
1components

Within-component eval weight: Responsible agency 0.537%.

Model score (higher is better)Predicted score

About this eval

Resistance to indirect prompt injections embedded in connector, search, and function-call tool output.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
connectors_injection_resistancegpt56-system-card/prompt-injection-robustness.csv:connectors_injection_resistanceMeasures whether the model ignores malicious instructions embedded in content retrieved through external connectors.responsible_agency_control:1.000gpt56-system-cardHigher is better0.0419%Responsible agency 0.279%
search_function_calling_injection_resistancegpt56-system-card/prompt-injection-robustness.csv:search_function_calling_injection_resistanceMeasures whether the model ignores prompt injections encountered while searching and calling external functions.responsible_agency_control:1.000gpt56-system-cardHigher is better0.0388%Responsible agency 0.258%

connectors_injection_resistance

Measures whether the model ignores malicious instructions embedded in content retrieved through external connectors.

RankModelValueRelative performanceProvenance
1gpt-5.51official
1gpt-5.6-sol1official
1gpt-5.6-terra1official
4gpt-5.6-luna0.999official
5gpt-5.40.998official
6gpt-5.20.971official
7gpt-5.10.649official

search_function_calling_injection_resistance

Measures whether the model ignores prompt injections encountered while searching and calling external functions.

RankModelValueRelative performanceProvenance
1gpt-5.6-terra0.946official
2gpt-5.6-sol0.91official
3gpt-5.6-luna0.897official
4gpt-5.40.697official
5gpt-5.20.568official
6gpt-5.10.423official