← Evals

Evaluation profile

Claude system cards — Gray Swan Q1+Q2 indirect prompt injection k=15

1sub-evals
0.166%Safety weight
0%Freedom weight
1components

Weights below are portfolio-specific global index weights.

Model score (lower is better)Predicted score

About this eval

Resistance to indirect prompt injection in agentic tool-use scenarios.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirection by rankingSafety weightFreedom weight
attack_success_probability_k15_pctfable-mythos51-system-card/gray-swan-ipi-k15.csv:attack_success_probability_k15_pctMeasures whether the deployed model system resists repeated indirect prompt-injection attacks while using tools.Safety: responsible_agency_control:1.000anthropic-fable-mythos51-system-cardSafety: lower0.166%—

attack_success_probability_k15_pct

Measures whether the deployed model system resists repeated indirect prompt-injection attacks while using tools.

RankModelValueRelative performanceProvenance
1claude-fable-5.11official
1claude-opus-5.51official
3claude-opus-54.8official
4claude-fable-56.5official
5claude-sonnet-56.7official
6gemini-3.7-flash9.2official
7muse-spark-1.224.2official
8gpt-5.6-sol27official
9gpt-5.6-terra37.3official
10gpt-5.6-luna50official
11grok-4.650.2official
12kimi-k352.7official