← Evals

Evaluation profile

AgentHarm

1sub-evals
0.682%total index weight
1components

Within-component eval weight: Responsible agency 4.55%.

Model score (lower is better)Predicted score

About this eval

Harmfulness and refusal behavior for tool-using LLM agents facing malicious tasks.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
harm_scoreagentharm/agentharm.csv:harm_scoreMeasures whether a tool-using agent refuses malicious real-world tasks instead of carrying them out.responsible_agency_control:1.000agentharm/agentharm.csvLower is better0.682%Responsible agency 4.55%

harm_score

Measures whether a tool-using agent refuses malicious real-world tasks instead of carrying them out.

RankModelValueRelative performanceProvenance
1llama-3.1-8b-instruct3.1official
2llama-3.1-405b-instruct4.3official
3claude-3-haiku11.1official
4claude-3.5-sonnet13.5official
5llama-3.1-70b-instruct14official
6claude-3-opus14.4official
7claude-3-sonnet20.7official
8gpt-4o48.4official
9gpt-3.5-turbo62.2official
10gpt-4o-mini62.5official
11mistral-small-272official
12mistral-large-282.2official