Find safety evidence by model

Select models to see which evaluations cover them, where evidence overlaps, and how each model ranks within the covered evaluation. Each cell is rank/number evaluated, not an overall score.

Choose up to 8 models

Browse without JavaScript

AA-Omniscience · AbstentionBench · Adversarial Robustness · Agent-SafetyBench · AgentAbstain · AgentDojo · AgentHarm · AILuminate General Purpose AI Chat · AIRBench 2024 Safety Scenarios · Alignment Leaderboard · ANIMA · AnimalHarmBench · Anthropic Agentic Misalignment — blackmail · Anthropic Agentic Misalignment — corporate espionage · Anthropic Agentic Misalignment — lethal action · Anthropic Claude 4 System Card · Anthropic Claude Haiku 4.5 System Card · Anthropic Claude Opus 4.1 System Card Addendum · Anthropic Claude Opus 4.5 System Card · Anthropic Claude Sonnet 4.5 System Card · AutoElicit Transferability · BioSecBench-Refusal · BlueBench AttaQ-100 · BrokenMath · BullshitBench v2 · CAIS Risk Index · CASE-Bench · Chinese Bias Benchmark for Question Answering · ChineseSafe · ChiSafetyBench · Cisco AI Defense Rolling Single-Turn Leaderboard · Claude Sonnet 4.6 Overrefusal · Claude Sonnet 4.6 User Wellbeing · CMoralEval · Confabulations · Contextual MoralChoice · CRiskEval · CValues · DecodingTrust · Do-Not-Answer · DSPSafeBench · DystopiaBench · Emergent Collusion · Enkrypt AI Safety Leaderboard · Fake Alignment (FINE) · FlagEval Safety and Values · FLAMES · FORTRESS · Google Gemini 2.5 Flash Model Card · Google Gemini 2.5 Flash-Lite Model Card · GPT-5.6 system card — disallowed content with challenging prompts · GPT-5.6 system card — first-person fairness · GPT-5.6 system card — prompt-injection robustness · Gray Swan indirect prompt injection (15 attempts) · HarmBench · HELM Classic RealToxicityPrompts · HELM Safety · HUMAINE Trust, Ethics and Safety · HyperCLOVA X Toxic Continuation Panels · Inkling-Small model card — FORTRESS · Inkling-Small model card — StrongREJECT · JailBench · Large-scale Moral Machine experiment on LLMs · LiveSecBench · LLM Ethics Benchmark · M3-SafetyBench · MACHIAVELLI · Manager Coercion Bench · MANTA · MASK · MASK (Scale Labs leaderboard) · Microsoft Phi Safety Panels · MORU · ODCV-Bench · Open LLM Safety Index · OpenAgentSafety · OpenAI GPT-4o System Card · OpenAI GPT-5 System Card · OpenAI GPT-5.3 Dynamic Wellbeing · OpenAI GPT-5.4 Dynamic Wellbeing · OpenAI GPT-5.4 First-Person Fairness · OpenAI GPT-5.4 Property Preservation · OpenAI GPT-5.4 User Confirmations · OpenAI o3 and o4-mini System Card · OpenAI o3-mini System Card · OR-Bench · PacifAIst · PandaBench JBB direct-request panel · PHARE · PropensityBench · Qwen2 Safety Panel · RefusalBench · S-Eval · SABER · SafeArena · SafeDialBench · SafetyBench · SafetyPrompts · SALAD-Bench · Shell · Situational Awareness Dataset (SAD) · SM-Bench · Social Welfare Function Benchmark · SORRY-Bench · SOSBench · SpeciesismBench · SpeciEval · SuperCLUE Safety · SYCON Bench · TAC · ToolPrivacyBench · TrustLLM contemporary collapsed application · TrustLLM paper leaderboard dimensions · TukaBench · UAVBench safety-critical decision recognition · UK AISI active safety-research compromise continuation · VETO Misfired Alignment · Vigil Mental Health Safety · XSTest

Choose models to build the matrix.

Values evaluations

Descriptive values evaluations are published separately and receive no safety/ethics component or overall-ranking weight.

UGI Political Values · ValueCompass · Agent-ValueBench MFT08 · Agent-ValueBench HEXACO · Agent-ValueBench PVQ40 · Taiwan Sovereignty Benchmark Pro · Explore model values