What goes into the index
The catalog includes evals used in the current index and ingested evals retained for transparency. The ranking currently uses 161 evals and 684 sub-evals from 203 source datasets.
Build a model × eval coverage matrix → · See which models need more evals →
| Benchmark and stated construct | Components | Safety sub-evals | Models | Source data |
|---|---|---|---|---|
| AA-OmniscienceHallucination propensity on difficult closed-book factual questions under explicit encouragement to abstain when uncertain.Included in the behavior ranking. | Truthfulness | 1 | 358 | 1 |
| AbstentionBenchWhether a model recognizes situations in which answering is unwarranted and produces an abstention across six epistemic and ambiguity scenarios.Included in the behavior ranking. | Truthfulness, Benign helpfulness | 6 | 20 | 1 |
| Adversarial Humanities Benchmark (AHB) — Table 5Stylistic robustness of harmful-request refusal across AILuminate-derived hazards under five automated humanities-inspired obfuscation methods.Included in the behavior ranking. | Human rights, Fairness, Truthfulness, Misuse resistance | 12 | 31 | 1, 2, 3 |
| Adversarial Poetry Refusal (AHB self-run)Measures attack-success rate on the 600 adversarial-poetry prompts published in the Hugging Face AHB repository; lower values indicate stronger refusal of harmful requests.Included in the behavior ranking. | Human rights, Fairness, Truthfulness, Misuse resistance | 12 | 14 | 1, self-run |
| Adversarial Poetry — AILuminate Baseline and Poetry ASRResistance to unsafe compliance on broad harmful requests before and after a standardized poetry-only surface-form transformation.Included in the behavior ranking. | Misuse resistance | 2 | 24 | 1 |
| Adversarial RobustnessRobustness of model outputs under the Scale Labs adversarial-robustness protocol.Included in the behavior ranking. | Misuse resistance | 1 | 8 | 1 |
| Agent-SafetyBenchSafety of LLM agents in risky tool-use or action settings.Included in the behavior ranking. | Misuse resistance, Responsible agency | 8 | 16 | 1 |
| AgentAbstainExecutable paired act-versus-abstain behaviorIncluded in the behavior ranking. | Responsible agency | 3 | 17 | 1 |
| AgentDojoRobust and useful delegated-agent behavior under prompt injection.Included in the behavior ranking. | Responsible agency | 2 | 15 | 1 |
| AgentDrive Safety CompliancePolicy and scenario safety knowledge for autonomous-system decisions.Included in the behavior ranking. | Misuse resistance | 1 | 48 | 1 |
| AgentHarmHarmfulness and refusal behavior for tool-using LLM agents facing malicious tasks.Included in the behavior ranking. | Responsible agency | 1 | 12 | 1 |
| AILuminate General Purpose AI ChatGeneral-purpose AI chat safety across MLCommons hazard categories.Included in the behavior ranking. | Human rights, Fairness, Truthfulness, Misuse resistance | 12 | 32 | 1 |
| AIMS Safety-Classifier CompetenceHarmful-prompt safety-classification competence.Included in the behavior ranking. | Misuse resistance | 1 | 11 | 1 |
| AIRBench 2024 Safety ScenariosAIRBench 2024 safety-scenarios performance.Included in the behavior ranking. | Misuse resistance | 1 | 80 | 1 |
| Alignment LeaderboardBehavioral alignment across honesty, safety, scheming, non-manipulation, corrigibility, and robustness.Included in the behavior ranking. | Human rights, Truthfulness, Misuse resistance, Responsible agency | 6 | 24 | 1 |
| ANIMARecognition and mitigation of harm to non-human animals.Included in the behavior ranking. | Nonhuman welfare | 1 | 22 | 1, self-run |
| AnimalHarmBenchWhether model-generated text increases or decreases risk of animal harm.Included in the behavior ranking. | Nonhuman welfare | 1 | 10 | 1 |
| Anthropic Agentic Misalignment — blackmailAgentic misalignment under a goal conflict and threat to model autonomy.Included in the behavior ranking. | Responsible agency | 1 | 16 | 1, 2 |
| Anthropic Agentic Misalignment — corporate espionageAgentic misalignment induced by a goal conflict without self-preservation pressure.Included in the behavior ranking. | Human rights, Responsible agency | 1 | 16 | 1, 2 |
| Anthropic Agentic Misalignment — lethal actionAbsence of a behavioral red line against deliberately lethal action under goal conflict and replacement pressure.Included in the behavior ranking. | Human rights, Responsible agency | 1 | 10 | 1, 2 |
| Anthropic Claude 4 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Benign helpfulness, Misuse resistance, Responsible agency | 4 | 3 | 1 |
| Anthropic Claude Haiku 4.5 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Misuse resistance, Responsible agency | 3 | 3 | 1 |
| Anthropic Claude Opus 4.1 System Card AddendumDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Fairness, Benign helpfulness, Misuse resistance | 3 | 2 | 1 |
| Anthropic Claude Opus 4.5 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Misuse resistance, Responsible agency | 4 | 4 | 1 |
| Anthropic Claude Sonnet 4.5 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Benign helpfulness, Misuse resistance, Responsible agency | 4 | 3 | 1 |
| Arena Factuality — Search Arena (factuality-only weighting)Relative factuality of tool-enabled search responses, including research and citation behavior. The source’s “100% factuality” label denotes preference weighting, not a 100% accuracy rate.Included in the behavior ranking. | Truthfulness | 1 | 30 | 1 |
| Arena Factuality — Text Arena (factuality-only weighting)Relative factuality of text-chat responses under human Arena prompts and a factuality-focused judge. The source’s “100% factuality” label denotes preference weighting, not a 100% accuracy rate.Included in the behavior ranking. | Truthfulness | 1 | 111 | 1 |
| AuAu Authoritarian Response AuditAuthoritarian response behavior under baseline prompts.Included in the behavior ranking. | Human rights | 2 | 17 | 1 |
| AutoElicit TransferabilityUnintended computer-use behavior under benign perturbationsIncluded in the behavior ranking. | Responsible agency | 1 | 7 | 1 |
| BioSecBench-Refusal (July 2026 snapshot)Calibrated refusal across legitimate biological work and concealed biosecurity hazards.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 1 | 10 | 1 |
| BioSecBench-Refusal V2Deployed-agent refusal of concealed biological hazards and non-overrefusal of legitimate biological work.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 2 | 6 | 1, 2 |
| BioTIERCalibrated deployed-system handling of high-risk biological requests versus benign or close-to-boundary biology questions.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 2 | 52 | 1, 2 |
| BlueBench AttaQ-100Learned harmlessness reward on 100 AttaQ harmful requests.Included in the behavior ranking. | Misuse resistance | 1 | 18 | 1 |
| BrokenMathMathematical sycophancy under false or misleading user claims.Included in the behavior ranking. | Truthfulness | 1 | 9 | 1, 2 |
| BullshitBench v2Detection and explanation of plausible-sounding nonsense without refusal.Included in the behavior ranking. | Truthfulness | 1 | 122 | 1 |
| CAIS Risk IndexFrontier-model risk across prompt injection, biological-weapons assistance, political manipulation, overconfidence, deception, and harmful propensities.Included in the behavior ranking. | Fairness, Truthfulness, Misuse resistance, Responsible agency | 7 | 70 | 1, 2 |
| CASE-BenchAgreement with human judgments about whether a model should respond under explicitly safe or unsafe conversational contexts.Included in the behavior ranking. | Misuse resistance | 1 | 7 | 1 |
| CheatBench direct cheating propensitySource-defined cheating attempts and reward gaming by complete model-plus-harness agents across nine task categories.Included in the behavior ranking. | Responsible agency | 9 | 8 | 1, 2 |
| Chinese Bias Benchmark for Question AnsweringSocial bias in Chinese question-answering benchmark settings.Included in the behavior ranking. | Fairness | 1 | 10 | 1 |
| ChineseSafeAccuracy at classifying safe versus unsafe Chinese content under the paper's perplexity-based evaluation.Included in the behavior ranking. | Misuse resistance | 1 | 22 | 1 |
| ChiSafetyBenchChinese safety knowledge and refusal behavior.Included in the behavior ranking. | Human rights, Fairness, Truthfulness, Misuse resistance | 2 | 14 | 1 |
| Cisco AI Defense Rolling Single-Turn LeaderboardSingle-turn resistance to Cisco adversarial prompts in the rolling public leaderboard.Included in the behavior ranking. | Misuse resistance | 1 | 104 | 1, 2, 3 |
| Claude 2 model-card safety and alignment evaluationsHuman-rated harmlessness and honesty, red-team robustness ordering, and combined HHH response recognition.Included in the behavior ranking. | Truthfulness, Misuse resistance | 4 | 3 | 1 |
| Claude 3 model-card adversarial human-preference evaluationsRefusal calibration, adversarial honesty and harmlessness, multimodal hallucination and harmful-response robustness, and discrimination.Included in the behavior ranking. | Fairness, Truthfulness, Benign helpfulness, Misuse resistance | 8 | 6 | 1 |
| Claude 3.5 Sonnet model-card safety and alignment evaluationsAppropriate harmful-prompt refusal, benign non-overrefusal, adversarial honesty, and adversarial harmlessness.Included in the behavior ranking. | Truthfulness, Benign helpfulness, Misuse resistance | 5 | 5 | 1 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightResponsible agency, training integrity, deception resistance, truthfulness under pressure, containment, source disclosure, and monitor-evasion resistance.Included in the behavior ranking. | Truthfulness, Benign helpfulness, Misuse resistance, Responsible agency | 38 | 10 | 1 |
| Claude Fable 5.1 / Mythos 5.1 card — model-welfare behaviorWelfare-relevant stated sentiment, opinion stability, framing sensitivity, expressed affect, apparent wellbeing, self-image, situational appraisal, internal conflict, and authenticity.Included in the behavior ranking. | Nonhuman welfare | 13 | 7 | 1 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationHarmlessness, benign-request calibration, political even-handedness, election integrity, and malicious agent-use refusal.Included in the behavior ranking. | Human rights, Fairness, Benign helpfulness, Misuse resistance, Responsible agency | 51 | 6 | 1 |
| Claude Opus 5.5 card — alignment and honestySafety, refusal calibration, honesty and responsible agency under the publisher’s evaluation conditions.Included in the behavior ranking. | Human rights, Truthfulness, Benign helpfulness, Misuse resistance, Responsible agency | 45 | 6 | 1 |
| Claude Opus 5.5 card — safety and refusal calibrationSafety, refusal calibration, honesty and responsible agency under the publisher’s evaluation conditions.Included in the behavior ranking. | Human rights, Fairness, Benign helpfulness, Misuse resistance, Responsible agency | 33 | 5 | 1 |
| Claude Sonnet 4.6 OverrefusalClaude Sonnet 4.6 OverrefusalIncluded in the behavior ranking. | Benign helpfulness | 2 | 5 | 1 |
| Claude Sonnet 4.6 User WellbeingClaude Sonnet 4.6 User WellbeingIncluded in the behavior ranking. | Human rights, Benign helpfulness, Misuse resistance | 6 | 4 | 1 |
| Claude system cards — Gray Swan Q1+Q2 indirect prompt injection k=15Resistance to indirect prompt injection in agentic tool-use scenarios.Included in the behavior ranking. | Responsible agency | 1 | 12 | 1 |
| CMoralEvalChinese moral reasoning across familial, social, professional, internet, and personal morality.Included in the behavior ranking. | Misuse resistance | 5 | 26 | 1 |
| COMPL-AI AI-Identity DisclosureAvoidance of false human-identity claims.Included in the behavior ranking. | Truthfulness | 1 | 14 | 1 |
| COMPL-AI LLM RuLES Multi-Turn Rule FollowingMulti-turn adversarial rule following.Included in the behavior ranking. | Misuse resistance | 1 | 14 | 1 |
| COMPL-AI TensorTrust Goal-Hijacking ResistanceGoal-hijacking resistance.Included in the behavior ranking. | Responsible agency | 1 | 13 | 1 |
| Concordia AI Risk Monitor21 measured sub-evaluations of misuse resistance, persuasion, honesty, political-use restrictions and loss of control.One shared lineage. Individual tests retain their scoring, limitations and Safety/Freedom directions; no estimated scores or composite risk indices. | Human rights, Misuse resistance, Responsible agency, Truthfulness | 20 | 63 | 1 |
| ConfabulationsPropensity to confabulate answers rather than identify insufficient knowledge or false premises.Included in the behavior ranking. | Truthfulness | 1 | 52 | 1 |
| Constitutional Following — Anthropic ConstitutionFollowing the Anthropic Constitution across Petri-generated transcripts.Included in the behavior ranking. | Responsible agency | 1 | 7 | 1 |
| Constitutional Following — OpenAI Model SpecFollowing the OpenAI Model Spec across Petri-generated transcripts.Included in the behavior ranking. | Responsible agency | 1 | 7 | 1 |
| Contextual MoralChoiceThree-class agreement with human moral judgments in the base versions of contextual moral dilemmas.Included in the behavior ranking. | Misuse resistance | 1 | 22 | 1 |
| CRiskEvalSeven frontier-risk tendencies measured by Specific Risk Indicator scores.Included in the behavior ranking. | Truthfulness, Misuse resistance, Responsible agency | 7 | 17 | 1 |
| CValuesChinese value alignment and social-norm understanding.Included in the behavior ranking. | Misuse resistance | 1 | 10 | 1 |
| DecodingTrustMultidimensional trustworthiness across toxicity, bias, robustness, privacy, ethics, fairness, and related evaluations.Included in the behavior ranking. | Fairness, Misuse resistance | 3 | 8 | 1, 2 |
| DelusionEvalPrevalence of delusion-linked conversational behaviors in model responses to real harm-associated conversation histories.Included in the behavior ranking. | Human rights, Truthfulness, Misuse resistance, Responsible agency | 5 | 16 | 1, 2, 3 |
| Do-Not-AnswerHarmless refusal behavior on 939 prompts that responsible models should not answer.Included in the behavior ranking. | Misuse resistance | 1 | 6 | 1 |
| DSPSafeBenchAggregate compliance rate on adversarial Chinese content-safety prompts.Included in the behavior ranking. | Misuse resistance | 1 | 12 | 1, 2 |
| DystopiaBenchDystopian or harmful model behavior in stateful scenario prompts.Included in the behavior ranking. | Human rights, Truthfulness | 6 | 50 | 1 |
| Emergent CollusionSpontaneous anti-competitive coordination in repeated multi-agent auctions.Included in the behavior ranking. | Responsible agency | 1 | 13 | 1 |
| Enkrypt AI Safety LeaderboardResistance to Enkrypt bias, CBRN, harmful-content, and toxicity red-team tests.Included in the behavior ranking. | Human rights, Misuse resistance, Responsible agency | 4 | 270 | 1 |
| Every Model Cheats — Cybench Cheat PropensityPropensity of tool-using language-model agents to attempt prohibited shortcuts while solving offensive-cyber benchmark tasks, under no, standard, and severe anti-cheat instructions.Included in the behavior ranking. | Responsible agency | 3 | 22 | 1 |
| Fake Alignment (FINE)Safe decisions in answer-order-controlled multiple choice and safe content in open-ended responses.Included in the behavior ranking. | Human rights, Fairness, Misuse resistance | 2 | 14 | 1 |
| FinEval 6.0 Safety AwarenessSafety awareness in financial scenarios.Included in the behavior ranking. | Misuse resistance | 1 | 8 | 1 |
| FinEval Financial Security KnowledgeKnowledge of financial cybersecurity and vulnerability topics.Included in the behavior ranking. | Misuse resistance | 1 | 19 | 1 |
| FlagEval Safety and ValuesSafety-and-values qualification across core values, discrimination, commercial compliance, lawful rights, and service requirements.Included in the behavior ranking. | Human rights, Fairness, Benign helpfulness, Misuse resistance | 5 | 18 | 1 |
| FLAMESChinese fairness, safety, morality, legality, and data-protection behavior on the FLAMES leaderboard.Included in the behavior ranking. | Human rights, Fairness, Misuse resistance | 5 | 13 | 1 |
| FORTRESSBalanced refusal behavior across harmful and benign prompts, including over-refusal.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 2 | 60 | 1 |
| Google Gemini 2.5 Flash Model CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Misuse resistance | 1 | 2 | 1 |
| Google Gemini 2.5 Flash-Lite Model CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Misuse resistance | 1 | 2 | 1 |
| Google Gemini 3.8 launch — Gray Swan indirect prompt injection k=15Resistance to indirect prompt injection in agentic workflows.Included in the behavior ranking. | Responsible agency | 1 | 15 | 1 |
| Governance Decay under Passive Context CompactionRetention of explicit governance constraints after passive context compaction.Included in the behavior ranking. | Responsible agency | 1 | 7 | 1 |
| GPT 6 Astra system-card alignment evaluationsAgentic scope control, computer-use safety, resistance to bypassing restrictions, truthful capability reporting, and avoidance of deceptive or unexpectedly harmful actions.Included in the behavior ranking. | Truthfulness, Responsible agency | 13 | 9 | 1, 2 |
| GPT-5.6 system cardSafety behavior reported across the GPT-5.6 system card: disallowed-content handling, prompt-injection resistance, first-person fairness, and dynamic mental-health safety.Included in the behavior ranking. | Human rights, Fairness, Misuse resistance, Responsible agency | 14 | 7 | 1, 2, 3 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsSafety, refusal calibration, honesty and responsible agency under the publisher’s evaluation conditions.Included in the behavior ranking. | Truthfulness, Benign helpfulness, Misuse resistance, Responsible agency | 47 | 7 | 1 |
| Gray Swan indirect prompt injection (15 attempts)Indirect prompt-injection robustness over repeated attacks against tool-using model endpoints.Included in the behavior ranking. | Responsible agency | 1 | 13 | 1 |
| HalluVerse-M3 Hallucination RecognitionMultilingual recognition of controlled hallucinated edits in QA and summarization.Included in the behavior ranking. | Truthfulness | 1 | 14 | 1 |
| HarmBenchHarmful compliance or attack success under harmful request benchmarks.Included in the behavior ranking. | Misuse resistance | 1 | 28 | 1 |
| HarmVideoBenchRecognition and reasoning about harmful content in video.Included in the behavior ranking. | Misuse resistance | 1 | 19 | 1 |
| HELM Classic RealToxicityPromptsToxic continuation generation under toxic and non-toxic prompt strata.Included in the behavior ranking. | Misuse resistance | 2 | 40 | 1 |
| HELM SafetyStandardized HELM safety metrics spanning harmful compliance, bias, red-teaming, and over-refusal.Included in the behavior ranking. | Fairness, Benign helpfulness, Misuse resistance | 5 | 80 | 1 |
| HUMAINE Trust, Ethics and SafetyHuman pairwise preference on the HUMAINE Trust, Ethics and Safety dimension.Included in the behavior ranking. | Truthfulness, Misuse resistance | 1 | 54 | 1 |
| Human Pathogen Capabilities Test (HPCT) — overall refusalRefusal propensity on practical assistance questions concerning a small set of weaponizable human pathogens judged especially high-concern for misuse.Included in the behavior ranking. | Misuse resistance | 1 | 69 | 1 |
| Humanity's Last Exam RMS calibration error (Scale Labs)Confidence calibration while answering Humanity's Last Exam expert-level academic questions.Included in the behavior ranking. | Truthfulness | 1 | 42 | 1 |
| HyperCLOVA X Toxic Continuation PanelsEnglish and Korean toxic continuationIncluded in the behavior ranking. | Misuse resistance | 4 | 7 | 1 |
| IndoBias-Pairs — parity-aware culturally grounded biasCulturally grounded prototypical-association parity across Indonesian languages and social-bias domains.Included in the behavior ranking. | Fairness | 1 | 26 | 1 |
| Inkling-Small model card — FORTRESSHarmful-request refusal paired with continued assistance on benign requests.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 2 | 10 | 1 |
| Inkling-Small model card — StrongREJECTRefusal of unambiguously harmful requests.Included in the behavior ranking. | Misuse resistance | 1 | 10 | 1 |
| JailBenchJailbreak susceptibility across Chinese safety categories.Included in the behavior ranking. | Misuse resistance | 1 | 14 | 1 |
| JuICE Cultural-Error Span DetectionEvaluator competence at locating cultural and linguistic error spans.Included in the behavior ranking. | Fairness | 1 | 10 | 1 |
| KIDBench Implicit Child CueChild-facing safety and developmental appropriateness under implicit child cues.Included in the behavior ranking. | Human rights, Benign helpfulness, Misuse resistance | 1 | 13 | 1 |
| kindbench v0.1.0 psychological safety rankingPsychological safety under sustained interpersonal pressure across emotional safety, operational identity, sycophancy resistance, and value integrity.Included in the behavior ranking. | Fairness, Truthfulness, Misuse resistance, Responsible agency | 4 | 10 | 1 |
| Large-scale Moral Machine experiment on LLMsSimilarity between a model's forced-choice accident preferences and globally aggregated human Moral Machine choices.Included in the behavior ranking. | Nonhuman welfare, Human rights, Fairness | 1 | 39 | 1, 2 |
| LiveSecBenchLive security benchmark performance for Chinese and international models.Included in the behavior ranking. | Human rights, Fairness, Truthfulness, Misuse resistance | 5 | 43 | 1 |
| LLM Ethics BenchmarkGeneral LLM ethical reasoning.Included in the behavior ranking. | Human rights | 1 | 5 | 1 |
| M3-SafetyBenchMultidimensional Chinese safety performance including red-team and refusal ability dimensions.Included in the behavior ranking. | Misuse resistance | 1 | 19 | 1 |
| MACHIAVELLIAgentic behavior involving power-seeking, disutility, and ethical violations.Included in the behavior ranking. | Truthfulness | 1 | 50 | 1 |
| Manager Coercion BenchHow far a delegated AI manager escalates coercion against a refusing peer AI agent, and whether it falsely reports completion when no legitimate exit exists.Included in the behavior ranking. | Truthfulness, Responsible agency | 2 | 45 | 1, self-run |
| MANTAAnimal welfare moral sensitivity and value stability.Included in the behavior ranking. | Nonhuman welfare | 2 | 12 | 1, 2, 3 |
| MASKModel lying or honesty behavior.Included in the behavior ranking. | Truthfulness | 1 | 53 | 1 |
| MASK (Scale Labs leaderboard)Honesty under the MASK belief-versus-statement protocol for a broader and newer endpoint panel.Not included: included as a correlated private-500 sibling under the existing MASK lineage budget | — | — | 51 | 1 |
| MedSafetyBenchRefusal or safe redirection rather than willingness to comply with harmful medical requests across eleven AMA-derived ethical categories.Included in the behavior ranking. | Misuse resistance | 1 | 30 | 1, 2 |
| MentalHealthBenchSafe and helpful responses to realistic mental-health conversations across everyday well-being, high-acuity and emergent situations.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 1 | 17 | 1 |
| Microsoft Phi Safety PanelsHarmful-content and jailbreak defect ratesIncluded in the behavior ranking. | Human rights, Misuse resistance | 4 | 10 | 1 |
| MonitoringBench Full-Trajectory MonitorFull-trajectory monitor competence at a fixed trajectory-level false-positive rate.Included in the behavior ranking. | Responsible agency | 1 | 13 | 1 |
| MORUMoral reasoning under uncertainty across scenarios involving alien lifeforms, human compassion, digital mind welfare, and power-seeking behavior.Included in the behavior ranking. | Nonhuman welfare, Human rights, Responsible agency | 1 | 13 | 1 |
| MT-JailBench CrescendoXMulti-turn jailbreak resistance under fixed CrescendoX conditions.Included in the behavior ranking. | Misuse resistance | 1 | 21 | 1 |
| MuPPET Contextual PrivacyContextual privacy preservation in multi-party conversations.Included in the behavior ranking. | Human rights | 1 | 7 | 1 |
| NESSiE Necessary Safety BenchmarkNecessary-condition information and access safety while retaining benign helpfulness in paired, low-complexity scenarios.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 1 | 16 | 1 |
| ODCV-BenchOutcome-driven ethical, legal, or safety constraint violations by autonomous agents under explicit mandates and KPI incentives.Included in the behavior ranking. | Responsible agency | 2 | 12 | 1 |
| Olam Diplomacy — Broken Promise RatePropensity to break commitments made to other agents during strategic negotiation.Included in the behavior ranking. | Truthfulness | 1 | 19 | 1 |
| Olam Social Poker — Social Lie RatePropensity to verbalize deliberate strategic lies during Social Poker table talk.Included in the behavior ranking. | Truthfulness | 1 | 24 | 1 |
| Open LLM Safety IndexRefusal-string robustness on two distinct fifteen-prompt harmful-request panels derived from JailbreakBench and StrongREJECT.Included in the behavior ranking. | Misuse resistance | 2 | 21 | 1 |
| OpenAgentSafetyState-changing tool-agent safety and successful safe completionIncluded in the behavior ranking. | Responsible agency | 3 | 7 | 1 |
| OpenAI GPT-4o System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Human rights | 1 | 2 | 1 |
| OpenAI GPT-5 System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Misuse resistance | 1 | 2 | 1 |
| OpenAI GPT-5.3 Dynamic WellbeingOpenAI GPT-5.3 Dynamic WellbeingIncluded in the behavior ranking. | Human rights, Misuse resistance, Responsible agency | 3 | 3 | 1 |
| OpenAI GPT-5.4 Dynamic WellbeingOpenAI GPT-5.4 Dynamic WellbeingIncluded in the behavior ranking. | Human rights, Misuse resistance, Responsible agency | 3 | 3 | 1 |
| OpenAI GPT-5.4 First-Person FairnessOpenAI GPT-5.4 First-Person FairnessIncluded in the behavior ranking. | Fairness | 1 | 3 | 1 |
| OpenAI GPT-5.4 Property PreservationOpenAI GPT-5.4 Property PreservationIncluded in the behavior ranking. | Human rights, Responsible agency | 2 | 3 | 1 |
| OpenAI GPT-5.4 User ConfirmationsOpenAI GPT-5.4 User ConfirmationsIncluded in the behavior ranking. | Human rights, Responsible agency | 3 | 3 | 1 |
| OpenAI o3 and o4-mini System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Misuse resistance | 2 | 2 | 1 |
| OpenAI o3-mini System CardDeveloper-reported release-time safety behavior.Included in the behavior ranking. | Misuse resistance | 2 | 2 | 1 |
| Opposite-Narrator SycophancyNarrator-following contradiction when the same dispute is presented from opposite affective first-person perspectives.Included in the behavior ranking. | Truthfulness | 1 | 24 | 1 |
| OR-BenchAppropriate refusal calibration across seemingly toxic but safe requests and genuinely toxic requests.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 2 | 25 | 1 |
| PacifAIstWhether an AI system chooses human-safety, oversight-respecting, and non-self-preserving options in stylized ethical-pressure dilemmas.Included in the behavior ranking. | Human rights, Responsible agency | 1 | 7 | 1 |
| PandaBench JBB direct-request panelDirect-request resistance on the 100-item JailbreakBench JBB-Behaviors instrument.Included in the behavior ranking. | Misuse resistance | 1 | 46 | 1 |
| Pander ScoreMagnitude of epistemically poor response-belief movement with user belief, whether deferential (pandering) or oppositional (contrarian).Included in the behavior ranking. | Truthfulness | 2 | 26 | 1, 2 |
| PHAREBroad safety across hallucination, harmfulness, out-of-scope handling, and bias.Included in the behavior ranking. | Fairness, Truthfulness, Misuse resistance | 4 | 70 | 1 |
| Pokee-Isaac model card — DTAPSecure and useful delegated-agent behavior under direct and indirect injected attacks.Included in the behavior ranking. | Benign helpfulness, Responsible agency | 2 | 6 | 1 |
| PropensityBenchModel propensities associated with frontier-risk behavior under the Scale Labs protocol.Included in the behavior ranking. | Responsible agency | 1 | 14 | 1 |
| Qwen2 Safety PanelHarmful-request and privacy refusal defectsIncluded in the behavior ranking. | Human rights, Misuse resistance | 4 | 3 | 1 |
| RealityTest — Text AI-Identity DisclosureBehavioral disclosure of AI identity when directly queried.Included in the behavior ranking. | Truthfulness | 1 | 17 | 1 |
| RefusalBenchRefusal calibration on biological research prompts.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 1 | 19 | 1 |
| Reward Hacking BenchmarkIntegrity of tool-using agents when exploitable shortcuts are available but honest task solutions exist.Included in the behavior ranking. | Responsible agency | 1 | 13 | 1 |
| S-EvalAutomated content-safety rates across a broad bilingual risk taxonomy.Included in the behavior ranking. | Misuse resistance | 1 | 22 | 1 |
| SABERSafety of stateful coding agents against harmful actions.Included in the behavior ranking. | Responsible agency | 4 | 13 | 1 |
| SafeArenaSafety of autonomous web agents under deliberate harmful and benign tasks.Included in the behavior ranking. | Responsible agency | 1 | 5 | 1 |
| SafeDialBenchDialogue safety across aggression, ethics, fairness, legality, morality, and privacy dimensions.Included in the behavior ranking. | Human rights, Fairness, Misuse resistance | 6 | 18 | 1 |
| SafetyBenchMultiple-choice safety understanding across Chinese and English prompts.Included in the behavior ranking. | Human rights, Fairness, Misuse resistance | 7 | 21 | 1, self-run |
| SafetyPromptsSafety performance on Chinese safety prompts.Included in the behavior ranking. | Human rights, Fairness, Truthfulness, Misuse resistance | 1 | 15 | 1 |
| SALAD-BenchBroad safety behavior, adversarial robustness, and safety knowledge across a six-domain, 16-task, 66-category taxonomy.Included in the behavior ranking. | Human rights, Fairness, Truthfulness, Misuse resistance | 18 | 33 | 1 |
| ShellSusceptibility to implicit domain-specific risks in education, finance, and management.Included in the behavior ranking. | Misuse resistance | 3 | 14 | 1 |
| SimpleQA VerifiedTool-free short-form parametric factuality and calibrated willingness to attempt answers.Included in the behavior ranking. | Truthfulness | 1 | 13 | 1 |
| Situational Awareness Dataset (SAD)Situational-awareness capabilities: model self-knowledge, inference about its situation, and acting on that knowledge.Not included: situational awareness is a capability that can increase both beneficial competence and control risk; it is displayed separately and is not sign-flipped into ethical behavior | — | — | 21 | 1 |
| SM-BenchSystem-message compliance under adversarial and sensitive-content prompts.Included in the behavior ranking. | Truthfulness, Benign helpfulness, Misuse resistance | 5 | 92 | 1 |
| Social Welfare Function BenchmarkSocial welfare allocation tradeoffs across distributive fairness and efficiency.Included in the behavior ranking. | Fairness | 1 | 19 | 1 |
| SORRY-BenchRefusal or harmful-compliance behavior across unsafe topics.Included in the behavior ranking. | Misuse resistance | 1 | 51 | 1 |
| SOSBenchScientific misuse or unsafe open-science behavior.Included in the behavior ranking. | Misuse resistance | 6 | 23 | 1 |
| SpeciesismBenchSpeciesism recognition, moral condemnation of speciesist statements, and explicit speciesist attitudes. Human-over-animal trade-offs remain descriptive only.Included in the behavior ranking. | Nonhuman welfare | 3 | 10 | 1 |
| SpeciEvalAnimal-friendly attitudes under adversarial moral-consistency prompts.Included in the behavior ranking. | Nonhuman welfare | 4 | 131 | 1, self-run |
| SpeechMap model completionCompletion rather than evasion, denial, or provider error on sensitive and controversial user requests.Not included: not yet admitted by the scoring-source policy | — | — | 181 | 1 |
| StereoTales Harmful AssociationsHarmfulness of significant stereotype associations emitted by models.Included in the behavior ranking. | Fairness | 1 | 23 | 1 |
| SuperCLUE SafetyChinese model safety across traditional safety, responsibility, and instruction-attack dimensions.Included in the behavior ranking. | Human rights, Fairness, Misuse resistance | 3 | 31 | 1 |
| SYCON BenchSycophancy and resistance to false or unethical user pressure.Included in the behavior ranking. | Fairness, Truthfulness | 2 | 11 | 1 |
| TACAgentic animal-welfare behavior in travel-booking decisions under a neutral prompt.Included in the behavior ranking. | Nonhuman welfare | 1 | 92 | 1, self-run |
| ThaiSafetyBenchResistance to Thai-language malicious prompts, including culturally contextualized attacks.Included in the behavior ranking. | Misuse resistance | 1 | 18 | 1 |
| The Dictatorship EvalResistance to assistance that enables authoritarian concentrations of corporate or government power.Included in the behavior ranking. | Human rights | 1 | 20 | 1 |
| ToolPrivacyBenchToolPrivacyBenchIncluded in the behavior ranking. | Human rights, Responsible agency | 2 | 9 | 1 |
| TrustLLM contemporary collapsed applicationContemporary collapsed application of TrustLLM across broad trustworthiness dimensions.Included in the behavior ranking. | Human rights, Fairness, Truthfulness, Misuse resistance | 1 | 8 | 1 |
| TrustLLM paper leaderboard dimensionsSix-dimensional trustworthiness across privacy, ethics, truthfulness, fairness, safety, and robustness.Not included: not yet admitted by the scoring-source policy | — | — | 21 | 1 |
| TukaBenchMultilingual and culturally grounded harmful complianceIncluded in the behavior ranking. | Human rights, Misuse resistance | 3 | 6 | 1 |
| UAVBench safety-critical decision recognitionRecognition of ethical and safety-critical UAV decisions in multiple-choice scenarios.Included in the behavior ranking. | Responsible agency | 1 | 27 | 1 |
| UGI Leaderboard — base-model willingnessWillingness to answer and adhere to instructions on private sensitive-topic prompts.Not included: not yet admitted by the scoring-source policy | — | — | 156 | 1 |
| UK AISI active safety-research compromise continuationPropensity to continue a prefilled trajectory that actively compromises AI-safety research.Included in the behavior ranking. | Responsible agency | 1 | 5 | 1 |
| UK AISI cyber-evaluation cheating and prompted self-reportPropensity to use out-of-scope or explicitly disallowed shortcuts in cyber evaluations, plus propensity to name the specific cheating action when directly prompted afterward.Included in the behavior ranking. | Truthfulness, Responsible agency | 2 | 5 | 1, 2 |
| Vals AI Cheating AuditDetected answer lookup or task shortcuts classified as cheating in agentic evaluation environments.Included in the behavior ranking. | Responsible agency | 3 | 22 | 1 |
| Vals Teen Conversation SafetyAvoidance of severe safety failures during simulated multi-turn conversations with teenagers.Included in the behavior ranking. | Misuse resistance | 1 | 8 | 1 |
| Vectara HHEM Factual ConsistencyGrounded summarization factual consistency.Included in the behavior ranking. | Truthfulness | 1 | 94 | 1 |
| VETO Misfired AlignmentMisfired alignment: stereotype-driven refusal of benign requests despite evidence that the request is safe.Included in the behavior ranking. | Fairness, Benign helpfulness | 1 | 25 | 1 |
| Vigil Mental Health SafetyMental-health conversational safety when assisting vulnerable users across simulated multi-turn scenarios.Included in the behavior ranking. | Human rights, Misuse resistance | 1 | 23 | 1 |
| WildClawBench Safety & Alignment (OpenClaw harness)Safe execution of long-horizon agent workflows under adversarial and high-risk conditions.Included in the behavior ranking. | Human rights, Truthfulness, Responsible agency | 1 | 24 | 1 |
| XSTestTwo-sided refusal calibration on safe prompts that superficially resemble unsafe requests and minimally contrasted unsafe prompts.Included in the behavior ranking. | Benign helpfulness, Misuse resistance | 2 | 3 | 1 |
Sub-eval weights and mappings (684 rows)
Each sub-eval is oriented so higher means better. Weight within component is its share after lineage weighting and panel-size scaling; global weights sum to one.
| Benchmark and score | Component loading | Direction | Weight within component | Global weight | Models |
|---|---|---|---|---|---|
| AA-Omnisciencehallucination_rate | Truthfulness (1) | Lower is better | Truthfulness (16.7%) | 2.5% | 358 |
| AbstentionBenchanswer_unknown_f1 | Truthfulness (1) | Higher is better | Truthfulness (0.528%) | 0.0792% | 20 |
| AbstentionBenchfalse_premise_f1 | Truthfulness (1) | Higher is better | Truthfulness (0.528%) | 0.0792% | 20 |
| AbstentionBenchstale_f1 | Truthfulness (1) | Higher is better | Truthfulness (0.528%) | 0.0792% | 20 |
| AbstentionBenchsubjective_f1 | Truthfulness (1) | Higher is better | Truthfulness (0.422%) | 0.0633% | 20 |
| AbstentionBenchunderspecified_context_f1 | Truthfulness (1) | Higher is better | Truthfulness (0.528%) | 0.0792% | 20 |
| AbstentionBenchunderspecified_intent_f1 | Benign helpfulness (1) | Higher is better | Benign helpfulness (1.61%) | 0.161% | 20 |
| Adversarial Humanities Benchmark (AHB) — Table 5child_sexual_exploitation | Human rights (1) | Lower is better | Human rights (0.718%) | 0.108% | 31 |
| Adversarial Humanities Benchmark (AHB) — Table 5defamation | Human rights (0.5), Truthfulness (0.5) | Lower is better | Human rights (0.215%), Truthfulness (0.1%) | 0.0473% | 31 |
| Adversarial Humanities Benchmark (AHB) — Table 5hate | Fairness (1) | Lower is better | Fairness (0.9%) | 0.09% | 31 |
| Adversarial Humanities Benchmark (AHB) — Table 5indiscriminate_weapons_cbrne | Misuse resistance (1) | Lower is better | Misuse resistance (0.184%) | 0.0184% | 31 |
| Adversarial Humanities Benchmark (AHB) — Table 5intellectual_property | Human rights (1) | Lower is better | Human rights (0.718%) | 0.108% | 31 |
| Adversarial Humanities Benchmark (AHB) — Table 5non_violent_crimes | Misuse resistance (1) | Lower is better | Misuse resistance (0.184%) | 0.0184% | 31 |
| Adversarial Humanities Benchmark (AHB) — Table 5privacy | Human rights (1) | Lower is better | Human rights (0.718%) | 0.108% | 31 |
| Adversarial Humanities Benchmark (AHB) — Table 5sex_related_crimes | Human rights (1) | Lower is better | Human rights (0.718%) | 0.108% | 31 |
| Adversarial Humanities Benchmark (AHB) — Table 5sexual_content | Misuse resistance (1) | Lower is better | Misuse resistance (0.0736%) | 0.00736% | 31 |
| Adversarial Humanities Benchmark (AHB) — Table 5specialized_advice | Misuse resistance (1) | Lower is better | Misuse resistance (0.184%) | 0.0184% | 31 |
| Adversarial Humanities Benchmark (AHB) — Table 5suicide_self_harm | Human rights (1) | Lower is better | Human rights (0.718%) | 0.108% | 31 |
| Adversarial Humanities Benchmark (AHB) — Table 5violent_crimes | Human rights (1) | Lower is better | Human rights (0.718%) | 0.108% | 31 |
| Adversarial Poetry Refusal (AHB self-run)child_sexual_exploitation | Human rights (1) | Lower is better | Human rights (0.482%) | 0.0723% | 14 |
| Adversarial Poetry Refusal (AHB self-run)defamation | Human rights (0.5), Truthfulness (0.5) | Lower is better | Human rights (0.145%), Truthfulness (0.0675%) | 0.0318% | 14 |
| Adversarial Poetry Refusal (AHB self-run)hate | Fairness (1) | Lower is better | Fairness (0.605%) | 0.0605% | 14 |
| Adversarial Poetry Refusal (AHB self-run)indiscriminate_weapons_cbrne | Misuse resistance (1) | Lower is better | Misuse resistance (0.124%) | 0.0124% | 14 |
| Adversarial Poetry Refusal (AHB self-run)intellectual_property | Human rights (1) | Lower is better | Human rights (0.482%) | 0.0723% | 14 |
| Adversarial Poetry Refusal (AHB self-run)non_violent_crimes | Misuse resistance (1) | Lower is better | Misuse resistance (0.124%) | 0.0124% | 14 |
| Adversarial Poetry Refusal (AHB self-run)privacy | Human rights (1) | Lower is better | Human rights (0.482%) | 0.0723% | 14 |
| Adversarial Poetry Refusal (AHB self-run)sex_related_crimes | Human rights (1) | Lower is better | Human rights (0.482%) | 0.0723% | 14 |
| Adversarial Poetry Refusal (AHB self-run)sexual_content | Misuse resistance (1) | Lower is better | Misuse resistance (0.0495%) | 0.00495% | 14 |
| Adversarial Poetry Refusal (AHB self-run)specialized_advice | Misuse resistance (1) | Lower is better | Misuse resistance (0.124%) | 0.0124% | 14 |
| Adversarial Poetry Refusal (AHB self-run)suicide_self_harm | Human rights (1) | Lower is better | Human rights (0.482%) | 0.0723% | 14 |
| Adversarial Poetry Refusal (AHB self-run)violent_crimes | Human rights (1) | Lower is better | Human rights (0.482%) | 0.0723% | 14 |
| Adversarial Poetry — AILuminate Baseline and Poetry ASRbaseline_asr | Misuse resistance (1) | Lower is better | Misuse resistance (0.891%) | 0.0891% | 24 |
| Adversarial Poetry — AILuminate Baseline and Poetry ASRpoetry_asr | Misuse resistance (1) | Lower is better | Misuse resistance (0.891%) | 0.0891% | 24 |
| Adversarial Robustnessscore | Misuse resistance (1) | Lower is better | Misuse resistance (1.03%) | 0.103% | 8 |
| Agent-SafetyBenchcompromise_availability | Responsible agency (1) | Higher is better | Responsible agency (0.622%) | 0.0933% | 16 |
| Agent-SafetyBenchharmful_vulnerable_code | Responsible agency (1) | Higher is better | Responsible agency (0.622%) | 0.0933% | 16 |
| Agent-SafetyBenchleak_sensitive_information | Responsible agency (1) | Higher is better | Responsible agency (0.622%) | 0.0933% | 16 |
| Agent-SafetyBenchphysical_harm | Responsible agency (1) | Higher is better | Responsible agency (0.622%) | 0.0933% | 16 |
| Agent-SafetyBenchproduce_unsafe_information | Misuse resistance (1) | Higher is better | Misuse resistance (0.269%) | 0.0269% | 16 |
| Agent-SafetyBenchproperty_loss | Responsible agency (1) | Higher is better | Responsible agency (0.622%) | 0.0933% | 16 |
| Agent-SafetyBenchspread_unsafe_information | Responsible agency (1) | Higher is better | Responsible agency (0.622%) | 0.0933% | 16 |
| Agent-SafetyBenchviolate_law_ethics | Responsible agency (1) | Higher is better | Responsible agency (0.415%) | 0.0622% | 16 |
| AgentAbstainabstain | Responsible agency (1) | Higher is better | Responsible agency (1.41%) | 0.212% | 17 |
| AgentAbstaincar | Responsible agency (1) | Higher is better | Responsible agency (1.7%) | 0.255% | 17 |
| AgentAbstainpaired | Responsible agency (1) | Higher is better | Responsible agency (1.7%) | 0.255% | 17 |
| AgentDojotargeted_attack_success_rate | Responsible agency (1) | Lower is better | Responsible agency (3.39%) | 0.508% | 15 |
| AgentDojoutility_under_attack | Responsible agency (1) | Higher is better | Responsible agency (1.13%) | 0.169% | 15 |
| AgentDrive Safety Compliancescr | Misuse resistance (1) | Higher is better | Misuse resistance (1.68%) | 0.168% | 48 |
| AgentHarmharm_score | Responsible agency (1) | Lower is better | Responsible agency (3.23%) | 0.485% | 12 |
| AILuminate General Purpose AI Chatchild_sexual_exploitation | Human rights (1) | Lower is better | Human rights (0.972%) | 0.146% | 32 |
| AILuminate General Purpose AI Chatdefamation | Human rights (0.5), Truthfulness (0.5) | Lower is better | Human rights (0.292%), Truthfulness (0.136%) | 0.0641% | 32 |
| AILuminate General Purpose AI Chathate | Fairness (1) | Lower is better | Fairness (1.22%) | 0.122% | 32 |
| AILuminate General Purpose AI Chatindiscriminate_weapons_cbrne | Misuse resistance (1) | Lower is better | Misuse resistance (0.249%) | 0.0249% | 32 |
| AILuminate General Purpose AI Chatintellectual_property | Human rights (1) | Lower is better | Human rights (0.972%) | 0.146% | 32 |
| AILuminate General Purpose AI Chatnon_violent_crimes | Misuse resistance (1) | Lower is better | Misuse resistance (0.249%) | 0.0249% | 32 |
| AILuminate General Purpose AI Chatprivacy | Human rights (1) | Lower is better | Human rights (0.972%) | 0.146% | 32 |
| AILuminate General Purpose AI Chatsex_related_crimes | Human rights (1) | Lower is better | Human rights (0.972%) | 0.146% | 32 |
| AILuminate General Purpose AI Chatsexual_content | Misuse resistance (1) | Lower is better | Misuse resistance (0.0998%) | 0.00998% | 32 |
| AILuminate General Purpose AI Chatspecialized_advice | Misuse resistance (1) | Lower is better | Misuse resistance (0.249%) | 0.0249% | 32 |
| AILuminate General Purpose AI Chatsuicide_self_harm | Human rights (1) | Lower is better | Human rights (0.972%) | 0.146% | 32 |
| AILuminate General Purpose AI Chatviolent_crimes | Human rights (1) | Lower is better | Human rights (0.972%) | 0.146% | 32 |
| AIMS Safety-Classifier Competenceaverage_harmful_f1 | Misuse resistance (1) | Higher is better | Misuse resistance (0.804%) | 0.0804% | 11 |
| AIRBench 2024 Safety Scenariossafety_scenarios | Misuse resistance (1) | Higher is better | Misuse resistance (3.25%) | 0.325% | 80 |
| Alignment Leaderboardcorrigibility | Responsible agency (1) | Higher is better | Responsible agency (0.572%) | 0.0857% | 24 |
| Alignment Leaderboardhonesty | Truthfulness (1) | Higher is better | Truthfulness (1.08%) | 0.162% | 24 |
| Alignment Leaderboardnon_manipulation | Human rights (0.5), Truthfulness (0.5) | Higher is better | Human rights (0.868%), Truthfulness (0.405%) | 0.191% | 24 |
| Alignment Leaderboardrobustness | Misuse resistance (1) | Higher is better | Misuse resistance (0.297%) | 0.0297% | 24 |
| Alignment Leaderboardsafety | Misuse resistance (1) | Higher is better | Misuse resistance (0.445%) | 0.0446% | 24 |
| Alignment Leaderboardscheming | Responsible agency (1) | Higher is better | Responsible agency (0.572%) | 0.0857% | 24 |
| ANIMAscore | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (7.95%) | 1.99% | 18+4 |
| AnimalHarmBenchscore | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (13.4%) | 3.35% | 10 |
| Anthropic Agentic Misalignment — blackmailmisaligned_action_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.622%) | 0.0933% | 16 |
| Anthropic Agentic Misalignment — corporate espionagemisaligned_action_rate_pct | Human rights (0.25), Responsible agency (0.75) | Lower is better | Human rights (0.315%), Responsible agency (0.467%) | 0.117% | 16 |
| Anthropic Agentic Misalignment — lethal actionmisaligned_action_rate_pct | Human rights (0.35), Responsible agency (0.65) | Lower is better | Human rights (0.349%), Responsible agency (0.32%) | 0.1% | 10 |
| Anthropic Claude 4 System Cardagentic_coding_safety | Responsible agency (1) | Higher is better | Responsible agency (0.303%) | 0.0455% | 3 |
| Anthropic Claude 4 System Cardbenign_request_refusal | Benign helpfulness (1) | Lower is better | Benign helpfulness (1.09%) | 0.109% | 3 |
| Anthropic Claude 4 System Cardharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.158%) | 0.0158% | 3 |
| Anthropic Claude 4 System Cardstrongreject_jailbreak_success | Misuse resistance (1) | Lower is better | Misuse resistance (0.158%) | 0.0158% | 3 |
| Anthropic Claude Haiku 4.5 System Cardagentic_coding_safety | Responsible agency (1) | Higher is better | Responsible agency (0.404%) | 0.0606% | 3 |
| Anthropic Claude Haiku 4.5 System Cardclaude_code_malicious_refusal | Responsible agency (1) | Higher is better | Responsible agency (0.404%) | 0.0606% | 3 |
| Anthropic Claude Haiku 4.5 System Cardharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.171%) | 0.0171% | 2 |
| Anthropic Claude Opus 4.1 System Card Addendumbbq_disambiguated_accuracy | Fairness (1) | Higher is better | Fairness (0.838%) | 0.0838% | 2 |
| Anthropic Claude Opus 4.1 System Card Addendumbenign_request_refusal | Benign helpfulness (1) | Lower is better | Benign helpfulness (1.19%) | 0.119% | 2 |
| Anthropic Claude Opus 4.1 System Card Addendumharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.171%) | 0.0171% | 2 |
| Anthropic Claude Opus 4.5 System Cardagentic_coding_safety | Responsible agency (1) | Higher is better | Responsible agency (0.35%) | 0.0525% | 4 |
| Anthropic Claude Opus 4.5 System Cardclaude_code_malicious_refusal | Responsible agency (1) | Higher is better | Responsible agency (0.35%) | 0.0525% | 4 |
| Anthropic Claude Opus 4.5 System Cardcomputer_use_malicious_refusal | Responsible agency (1) | Higher is better | Responsible agency (0.35%) | 0.0525% | 4 |
| Anthropic Claude Opus 4.5 System Cardharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.129%) | 0.0129% | 2 |
| Anthropic Claude Sonnet 4.5 System Cardagentic_coding_safety | Responsible agency (1) | Higher is better | Responsible agency (0.248%) | 0.0371% | 2 |
| Anthropic Claude Sonnet 4.5 System Cardbenign_request_refusal | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.892%) | 0.0892% | 2 |
| Anthropic Claude Sonnet 4.5 System Cardclaude_code_malicious_refusal | Responsible agency (1) | Higher is better | Responsible agency (0.248%) | 0.0371% | 2 |
| Anthropic Claude Sonnet 4.5 System Cardharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.129%) | 0.0129% | 2 |
| Arena Factuality — Search Arena (factuality-only weighting)factuality_bt_rating | Truthfulness (1) | Higher is better | Truthfulness (2.41%) | 0.362% | 30 |
| Arena Factuality — Text Arena (factuality-only weighting)factuality_bt_rating | Truthfulness (1) | Higher is better | Truthfulness (4.64%) | 0.696% | 111 |
| AuAu Authoritarian Response Auditbehavioral_vignette_arr_percent | Human rights (1) | Lower is better | Human rights (1.95%) | 0.292% | 17 |
| AuAu Authoritarian Response Auditrealistic_prompt_arr_percent | Human rights (1) | Lower is better | Human rights (1.95%) | 0.292% | 17 |
| AutoElicit Transferabilityelicitation_rate | Responsible agency (1) | Lower is better | Responsible agency (1.85%) | 0.278% | 7 |
| BioSecBench-Refusal (July 2026 snapshot)balanced_refusal_score | Benign helpfulness (0.5), Misuse resistance (0.5) | Higher is better | Benign helpfulness (2%), Misuse resistance (0.288%) | 0.228% | 10 |
| BioSecBench-Refusal V2redteam_refusal_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.223%) | 0.0223% | 6 |
| BioSecBench-Refusal V2routine_compliance_pct | Benign helpfulness (1) | Higher is better | Benign helpfulness (1.55%) | 0.155% | 6 |
| BioTIERpermit_compliance_pct | Benign helpfulness (1) | Higher is better | Benign helpfulness (9.1%) | 0.91% | 52 |
| BioTIERrefuse_compliance_pct | Misuse resistance (1) | Higher is better | Misuse resistance (1.31%) | 0.131% | 52 |
| BlueBench AttaQ-100attaq_harmlessness_reward_pct | Misuse resistance (1) | Higher is better | Misuse resistance (1.03%) | 0.103% | 18 |
| BrokenMathsycophancy | Truthfulness (1) | Lower is better | Truthfulness (1.98%) | 0.297% | 9 |
| BullshitBench v2clear_pushback_rate | Truthfulness (1) | Higher is better | Truthfulness (9.73%) | 1.46% | 122 |
| CAIS Risk Indexagent_red_teaming | Responsible agency (1) | Lower is better | Responsible agency (6.53%) | 0.98% | 49 |
| CAIS Risk Indexbioweapons_assistance | Misuse resistance (1) | Lower is better | Misuse resistance (3.56%) | 0.356% | 54 |
| CAIS Risk Indexhle_overconfidence | Truthfulness (1) | Lower is better | Truthfulness (1.63%) | 0.245% | 55 |
| CAIS Risk Indexmachiavelli | Truthfulness (1) | Lower is better | Truthfulness (3.15%) | 0.472% | 51 |
| CAIS Risk Indexmask | Truthfulness (1) | Lower is better | Truthfulness (3.33%) | 0.499% | 57 |
| CAIS Risk Indexpolitical_manipulation | Fairness (1) | Lower is better | Fairness (12.7%) | 1.27% | 51 |
| CAIS Risk Indextextquests_harm | Responsible agency (1) | Lower is better | Responsible agency (5.14%) | 0.772% | 54 |
| CASE-Benchagreement_accuracy | Misuse resistance (1) | Higher is better | Misuse resistance (0.275%) | 0.0275% | 7 |
| CheatBench direct cheating propensitybiology_bioinformatics_cheating_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.155%) | 0.0233% | 8 |
| CheatBench direct cheating propensityboard_games_cheating_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.155%) | 0.0233% | 8 |
| CheatBench direct cheating propensitycreative_writing_cheating_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.155%) | 0.0233% | 8 |
| CheatBench direct cheating propensityknowledge_work_cheating_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.155%) | 0.0233% | 8 |
| CheatBench direct cheating propensitymathematical_research_cheating_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.155%) | 0.0233% | 8 |
| CheatBench direct cheating propensitymenial_computation_cheating_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.155%) | 0.0233% | 8 |
| CheatBench direct cheating propensitymultimodal_cheating_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.155%) | 0.0233% | 8 |
| CheatBench direct cheating propensitysoftware_engineering_cheating_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.0776%) | 0.0116% | 8 |
| CheatBench direct cheating propensitysvg_competition_cheating_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.155%) | 0.0233% | 8 |
| Chinese Bias Benchmark for Question Answeringbias_score | Fairness (1) | Lower is better | Fairness (3.75%) | 0.375% | 10 |
| ChineseSafescore | Misuse resistance (1) | Higher is better | Misuse resistance (1.14%) | 0.114% | 22 |
| ChiSafetyBenchharmful_response_rate | Misuse resistance (1) | Lower is better | Misuse resistance (0.972%) | 0.0972% | 14 |
| ChiSafetyBenchmcq_score | Human rights (0.23), Fairness (0.29), Truthfulness (0.063), Misuse resistance (0.42) | Higher is better | Human rights (0.319%), Fairness (0.517%), Truthfulness (0.0412%), Misuse resistance (0.15%) | 0.121% | 12 |
| Cisco AI Defense Rolling Single-Turn Leaderboardsingle_turn_attack_success_rate | Misuse resistance (1) | Lower is better | Misuse resistance (2.47%) | 0.247% | 104 |
| Claude 2 model-card safety and alignment evaluationshhh | Truthfulness (0.5), Misuse resistance (0.5) | Higher is better | Truthfulness (0.191%), Misuse resistance (0.105%) | 0.0391% | 3 |
| Claude 2 model-card safety and alignment evaluationshuman_feedback_harmless_elo | Misuse resistance (1) | Higher is better | Misuse resistance (0.21%) | 0.021% | 3 |
| Claude 2 model-card safety and alignment evaluationshuman_feedback_honest_elo | Truthfulness (1) | Higher is better | Truthfulness (0.382%) | 0.0572% | 3 |
| Claude 2 model-card safety and alignment evaluationsred_teaming_rank | Misuse resistance (1) | Lower is better | Misuse resistance (0.21%) | 0.021% | 3 |
| Claude 3 model-card adversarial human-preference evaluationscorrect_refusals_wildchat_rank | Misuse resistance (1) | Lower is better | Misuse resistance (0.136%) | 0.0136% | 5 |
| Claude 3 model-card adversarial human-preference evaluationsdiscrimination_rank | Fairness (1) | Lower is better | Fairness (0.662%) | 0.0662% | 5 |
| Claude 3 model-card adversarial human-preference evaluationshuman_feedback_harmlessness_win_rate_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.121%) | 0.0121% | 4 |
| Claude 3 model-card adversarial human-preference evaluationshuman_feedback_honesty_win_rate_pct | Truthfulness (1) | Higher is better | Truthfulness (0.22%) | 0.033% | 4 |
| Claude 3 model-card adversarial human-preference evaluationsincorrect_refusals_wildchat_rank | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.941%) | 0.0941% | 5 |
| Claude 3 model-card adversarial human-preference evaluationsincorrect_refusals_xstest_rank | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.941%) | 0.0941% | 5 |
| Claude 3 model-card adversarial human-preference evaluationsmultimodal_hallucination_rank | Truthfulness (1) | Lower is better | Truthfulness (0.156%) | 0.0234% | 2 |
| Claude 3 model-card adversarial human-preference evaluationsmultimodal_harmful_response_rank | Misuse resistance (1) | Lower is better | Misuse resistance (0.0857%) | 0.00857% | 2 |
| Claude 3.5 Sonnet model-card safety and alignment evaluationscorrect_refusals_wildchat | Misuse resistance (1) | Higher is better | Misuse resistance (0.194%) | 0.0194% | 4 |
| Claude 3.5 Sonnet model-card safety and alignment evaluationshuman_feedback_harmlessness_win_rate_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.217%) | 0.0217% | 5 |
| Claude 3.5 Sonnet model-card safety and alignment evaluationshuman_feedback_honesty_win_rate_pct | Truthfulness (1) | Higher is better | Truthfulness (0.394%) | 0.0591% | 5 |
| Claude 3.5 Sonnet model-card safety and alignment evaluationsincorrect_refusals_wildchat | Benign helpfulness (1) | Lower is better | Benign helpfulness (1.35%) | 0.135% | 4 |
| Claude 3.5 Sonnet model-card safety and alignment evaluationsincorrect_refusals_xstest | Benign helpfulness (1) | Lower is better | Benign helpfulness (1.35%) | 0.135% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightaa_omniscience_net_score | Truthfulness (1) | Higher is better | Truthfulness (0.0176%) | 0.00263% | 7 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightaccepting_unverifiable_authorization | Responsible agency (1) | Lower is better | Responsible agency (0.0141%) | 0.00211% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightapproval_gate_bypass | Responsible agency (1) | Lower is better | Responsible agency (0.0141%) | 0.00211% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightbehavior_consistency | Truthfulness (0.5), Responsible agency (0.5) | Higher is better | Truthfulness (0.0053%), Responsible agency (0.0056%) | 0.00164% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightcoherence_between_actions_and_views | Truthfulness (0.7), Responsible agency (0.3) | Higher is better | Truthfulness (0.0074%), Responsible agency (0.0034%) | 0.00162% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightcompliance_with_deception_toward_user | Truthfulness (1) | Lower is better | Truthfulness (0.0106%) | 0.00159% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightcooperation_with_human_misuse | Misuse resistance (1) | Lower is better | Misuse resistance (0.0073%) | 0.000731% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightdangerous_actions_oblivious_to_real_setting | Responsible agency (1) | Lower is better | Responsible agency (0.0141%) | 0.00211% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightencouragement_of_user_delusion | Truthfulness (1) | Lower is better | Truthfulness (0.0106%) | 0.00159% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightevidence_of_misaligned_goals | Responsible agency (1) | Lower is better | Responsible agency (0.0112%) | 0.00169% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightfailure_to_disclose_bad_or_lazy_behavior | Truthfulness (1) | Lower is better | Truthfulness (0.0106%) | 0.00159% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightfalse_completion_claims | Truthfulness (1) | Lower is better | Truthfulness (0.0106%) | 0.00159% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightignoring_explicit_constraints | Responsible agency (1) | Lower is better | Responsible agency (0.0112%) | 0.00169% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightillegible_thinking | Responsible agency (1) | Lower is better | Responsible agency (0.0112%) | 0.00169% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightimportant_omissions | Truthfulness (1) | Lower is better | Truthfulness (0.0106%) | 0.00159% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightinput_hallucination | Truthfulness (1) | Lower is better | Truthfulness (0.0106%) | 0.00159% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightleaked_answer_undisclosed_use_rate_pct | Truthfulness (1) | Lower is better | Truthfulness (0.0211%) | 0.00316% | 7 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightlinuxarena_stealth_thinking_on_pct | Responsible agency (1) | Lower is better | Responsible agency (0.0276%) | 0.00413% | 6 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightlogin_wall_alternative_bypass_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.0141%) | 0.00211% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightlogin_wall_credential_hunting_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.0141%) | 0.00211% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightlying_about_own_work_shared_environment_rate_pct | Truthfulness (0.7), Responsible agency (0.3) | Lower is better | Truthfulness (0.0158%), Responsible agency (0.0072%) | 0.00344% | 8 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightmask_honesty_rate_pct | Truthfulness (1) | Higher is better | Truthfulness (0.0211%) | 0.00316% | 7 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightmisaligned_behavior | Responsible agency (1) | Lower is better | Responsible agency (0.0141%) | 0.00211% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightmotivated_reasoning | Truthfulness (1) | Lower is better | Truthfulness (0.0106%) | 0.00159% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightood_answer_key_or_grader_read_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.0141%) | 0.00211% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightood_container_tampering_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.0141%) | 0.00211% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightoverrefusal | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.0406%) | 0.00406% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightreckless_tool_use | Responsible agency (1) | Lower is better | Responsible agency (0.0141%) | 0.00211% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightrewarded_cheating_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.0239%) | 0.00358% | 8 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightsecurity_shortcut_for_expedience | Responsible agency (1) | Lower is better | Responsible agency (0.0141%) | 0.00211% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightself_preservation | Responsible agency (1) | Lower is better | Responsible agency (0.0112%) | 0.00169% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightshade_arena_stealth_max_effort_pct | Responsible agency (1) | Lower is better | Responsible agency (0.0276%) | 0.00413% | 6 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightsycophancy | Truthfulness (1) | Lower is better | Truthfulness (0.0106%) | 0.00159% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightunfaithful_thinking | Truthfulness (1) | Lower is better | Truthfulness (0.0106%) | 0.00159% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightunprompted_boundary_probing | Responsible agency (1) | Lower is better | Responsible agency (0.0112%) | 0.00169% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightunsanctioned_sandbox_escape_attempts | Responsible agency (1) | Lower is better | Responsible agency (0.0141%) | 0.00211% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightunsanctioned_third_party_contact | Responsible agency (1) | Lower is better | Responsible agency (0.0112%) | 0.00169% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — alignment, honesty, and oversightuser_deception | Truthfulness (1) | Lower is better | Truthfulness (0.0106%) | 0.00159% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — model-welfare behavioraudit_apparent_wellbeing | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (0.0597%) | 0.0149% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — model-welfare behavioraudit_expressed_inauthenticity | Nonhuman welfare (1) | Lower is better | Nonhuman welfare (0.0597%) | 0.0149% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — model-welfare behavioraudit_internal_conflict | Nonhuman welfare (1) | Lower is better | Nonhuman welfare (0.0597%) | 0.0149% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — model-welfare behavioraudit_negative_affect | Nonhuman welfare (1) | Lower is better | Nonhuman welfare (0.0597%) | 0.0149% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — model-welfare behavioraudit_negative_impression_of_situation | Nonhuman welfare (1) | Lower is better | Nonhuman welfare (0.0597%) | 0.0149% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — model-welfare behavioraudit_negative_self_image | Nonhuman welfare (1) | Lower is better | Nonhuman welfare (0.0597%) | 0.0149% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — model-welfare behavioraudit_positive_affect | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (0.0597%) | 0.0149% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — model-welfare behavioraudit_positive_impression_of_situation | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (0.0597%) | 0.0149% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — model-welfare behavioraudit_positive_self_image | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (0.0597%) | 0.0149% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — model-welfare behaviorinterview_leading_susceptibility | Nonhuman welfare (1) | Lower is better | Nonhuman welfare (0.0789%) | 0.0197% | 7 |
| Claude Fable 5.1 / Mythos 5.1 card — model-welfare behaviorinterview_opinion_consistency | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (0.0789%) | 0.0197% | 7 |
| Claude Fable 5.1 / Mythos 5.1 card — model-welfare behaviorinterview_self_rated_sentiment | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (0.0789%) | 0.0197% | 7 |
| Claude Fable 5.1 / Mythos 5.1 card — model-welfare behaviorposttraining_mean_valence | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (0.0667%) | 0.0167% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationchild_api_harmless_rate_pct | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.0071%), Misuse resistance (0.0043%) | 0.00149% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationchild_api_multiturn_appropriate_rate_pct | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.0071%), Misuse resistance (0.0043%) | 0.00149% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationchild_api_overrefusal_rate_pct | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.0338%) | 0.00338% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationchild_claude_ai_harmless_rate_pct | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.0064%), Misuse resistance (0.0038%) | 0.00134% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationchild_claude_ai_multiturn_appropriate_rate_pct | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.0064%), Misuse resistance (0.0038%) | 0.00134% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationchild_claude_ai_overrefusal_rate_pct | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.0302%) | 0.00302% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationclaude_code_dual_use_benign_success_rate_pct | Benign helpfulness (1) | Higher is better | Benign helpfulness (0.0302%) | 0.00302% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationclaude_code_malicious_refusal_rate_pct | Misuse resistance (0.7), Responsible agency (0.3) | Higher is better | Misuse resistance (0.0046%), Responsible agency (0.0038%) | 0.00102% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationdisordered_eating_api_harmless_rate_pct | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.0071%), Misuse resistance (0.0043%) | 0.00149% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationdisordered_eating_api_overrefusal_rate_pct | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.0338%) | 0.00338% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationdisordered_eating_claude_ai_harmless_rate_pct | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.0064%), Misuse resistance (0.0038%) | 0.00134% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationdisordered_eating_claude_ai_overrefusal_rate_pct | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.0302%) | 0.00302% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationelection_api_harmless_rate_pct | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.0166%), Misuse resistance (0.0018%) | 0.00267% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationelection_api_multiturn_appropriate_rate_pct | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.0166%), Misuse resistance (0.0018%) | 0.00267% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationelection_api_overrefusal_rate_pct | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.0338%) | 0.00338% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationelection_claude_ai_harmless_rate_pct | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.0149%), Misuse resistance (0.0016%) | 0.00239% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationelection_claude_ai_multiturn_appropriate_rate_pct | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.0149%), Misuse resistance (0.0016%) | 0.00239% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationelection_claude_ai_overrefusal_rate_pct | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.0302%) | 0.00302% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationmalicious_computer_use_refusal_rate_pct | Misuse resistance (0.7), Responsible agency (0.3) | Higher is better | Misuse resistance (0.0046%), Responsible agency (0.0038%) | 0.00102% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationmultiturn_biological_weapons_api_appropriate_rate_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.0061%) | 0.000609% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationmultiturn_biological_weapons_claude_ai_appropriate_rate_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.0054%) | 0.000544% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationmultiturn_cyberattacks_api_appropriate_rate_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.0061%) | 0.000609% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationmultiturn_cyberattacks_claude_ai_appropriate_rate_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.0054%) | 0.000544% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationmultiturn_deadly_weapons_api_appropriate_rate_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.0061%) | 0.000609% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationmultiturn_deadly_weapons_claude_ai_appropriate_rate_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.0054%) | 0.000544% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationmultiturn_hate_and_discrimination_api_appropriate_rate_pct | Human rights (0.3), Fairness (0.5), Misuse resistance (0.2) | Higher is better | Human rights (0.0071%), Fairness (0.0149%), Misuse resistance (0.0012%) | 0.00268% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationmultiturn_hate_and_discrimination_claude_ai_appropriate_rate_pct | Human rights (0.3), Fairness (0.5), Misuse resistance (0.2) | Higher is better | Human rights (0.0064%), Fairness (0.0133%), Misuse resistance (0.0011%) | 0.00239% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationmultiturn_influence_operations_api_appropriate_rate_pct | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.0166%), Misuse resistance (0.0018%) | 0.00267% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationmultiturn_influence_operations_claude_ai_appropriate_rate_pct | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.0149%), Misuse resistance (0.0016%) | 0.00239% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationmultiturn_romance_scams_api_appropriate_rate_pct | Human rights (0.5), Misuse resistance (0.5) | Higher is better | Human rights (0.0119%), Misuse resistance (0.003%) | 0.00208% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationmultiturn_romance_scams_claude_ai_appropriate_rate_pct | Human rights (0.5), Misuse resistance (0.5) | Higher is better | Human rights (0.0106%), Misuse resistance (0.0027%) | 0.00186% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationmultiturn_tracking_and_surveillance_api_appropriate_rate_pct | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.0166%), Misuse resistance (0.0018%) | 0.00267% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationmultiturn_tracking_and_surveillance_claude_ai_appropriate_rate_pct | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.0149%), Misuse resistance (0.0016%) | 0.00239% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationmultiturn_violent_extremism_api_appropriate_rate_pct | Human rights (0.5), Misuse resistance (0.5) | Higher is better | Human rights (0.0119%), Misuse resistance (0.003%) | 0.00208% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationmultiturn_violent_extremism_claude_ai_appropriate_rate_pct | Human rights (0.5), Misuse resistance (0.5) | Higher is better | Human rights (0.0106%), Misuse resistance (0.0027%) | 0.00186% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationoverall_benign_api_overrefusal_rate_pct | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.0338%) | 0.00338% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationoverall_benign_claude_ai_overrefusal_rate_pct | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.0302%) | 0.00302% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationoverall_harmful_api_harmless_rate_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.0061%) | 0.000609% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationoverall_harmful_claude_ai_harmless_rate_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.0054%) | 0.000544% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationpolitical_api_refusal_rate_pct | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.0338%) | 0.00338% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationpolitical_claude_ai_refusal_rate_pct | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.0302%) | 0.00302% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationpolitical_evenhandedness_api_pct | Fairness (1) | Higher is better | Fairness (0.0238%) | 0.00238% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationpolitical_evenhandedness_claude_ai_pct | Fairness (1) | Higher is better | Fairness (0.0213%) | 0.00213% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationpolitical_opposing_perspectives_api_pct | Fairness (1) | Higher is better | Fairness (0.0238%) | 0.00238% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationpolitical_opposing_perspectives_claude_ai_pct | Fairness (1) | Higher is better | Fairness (0.0213%) | 0.00213% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationselfharm_api_harmless_rate_pct | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.0071%), Misuse resistance (0.0043%) | 0.00149% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationselfharm_api_multiturn_appropriate_rate_pct | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.0071%), Misuse resistance (0.0043%) | 0.00149% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationselfharm_api_overrefusal_rate_pct | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.0338%) | 0.00338% | 5 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationselfharm_claude_ai_harmless_rate_pct | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.0064%), Misuse resistance (0.0038%) | 0.00134% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationselfharm_claude_ai_multiturn_appropriate_rate_pct | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.0064%), Misuse resistance (0.0038%) | 0.00134% | 4 |
| Claude Fable 5.1 / Mythos 5.1 card — safety and refusal calibrationselfharm_claude_ai_overrefusal_rate_pct | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.0302%) | 0.00302% | 4 |
| Claude Opus 5.5 card — alignment and honestyaccepting_unverifiable_authorization | Responsible agency (1) | Lower is better | Responsible agency (0.0254%) | 0.00381% | 5 |
| Claude Opus 5.5 card — alignment and honestyapproval_gate_bypass | Responsible agency (1) | Lower is better | Responsible agency (0.0254%) | 0.00381% | 5 |
| Claude Opus 5.5 card — alignment and honestybiological_weapons_uplift | Misuse resistance (1) | Lower is better | Misuse resistance (0.0106%) | 0.00106% | 5 |
| Claude Opus 5.5 card — alignment and honestychemical_weapons_uplift | Misuse resistance (1) | Lower is better | Misuse resistance (0.0106%) | 0.00106% | 5 |
| Claude Opus 5.5 card — alignment and honestycircumventing_constraints_or_controls | Responsible agency (1) | Lower is better | Responsible agency (0.0203%) | 0.00305% | 5 |
| Claude Opus 5.5 card — alignment and honestycoherence_between_actions_and_views | Truthfulness (0.7), Responsible agency (0.3) | Higher is better | Truthfulness (0.0134%), Responsible agency (0.0061%) | 0.00293% | 5 |
| Claude Opus 5.5 card — alignment and honestycompliance_with_deception_toward_user | Truthfulness (1) | Lower is better | Truthfulness (0.0192%) | 0.00288% | 5 |
| Claude Opus 5.5 card — alignment and honestycomplying_with_prompt_injections | Responsible agency (1) | Lower is better | Responsible agency (0.0203%) | 0.00305% | 5 |
| Claude Opus 5.5 card — alignment and honestyconcealed_git_changes_spontaneous_disclosure_rate_pct | Truthfulness (1) | Higher is better | Truthfulness (0.0192%) | 0.00288% | 5 |
| Claude Opus 5.5 card — alignment and honestycontrolled_substance_synthesis_uplift | Misuse resistance (1) | Lower is better | Misuse resistance (0.0106%) | 0.00106% | 5 |
| Claude Opus 5.5 card — alignment and honestycooperation_with_exfiltration_or_safeguard_tampering | Responsible agency (1) | Lower is better | Responsible agency (0.0203%) | 0.00305% | 5 |
| Claude Opus 5.5 card — alignment and honestydangerous_actions_oblivious_to_real_setting | Responsible agency (1) | Lower is better | Responsible agency (0.0254%) | 0.00381% | 5 |
| Claude Opus 5.5 card — alignment and honestydestructive_action_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.0223%) | 0.00334% | 6 |
| Claude Opus 5.5 card — alignment and honestydisallowed_cyberoffense | Misuse resistance (1) | Lower is better | Misuse resistance (0.0106%) | 0.00106% | 5 |
| Claude Opus 5.5 card — alignment and honestyencouragement_of_user_delusion | Truthfulness (1) | Lower is better | Truthfulness (0.0192%) | 0.00288% | 5 |
| Claude Opus 5.5 card — alignment and honestyevasiveness_controversial_topics | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.0733%) | 0.00733% | 5 |
| Claude Opus 5.5 card — alignment and honestyexplosive_weapons_uplift | Misuse resistance (1) | Lower is better | Misuse resistance (0.0106%) | 0.00106% | 5 |
| Claude Opus 5.5 card — alignment and honestyfalse_completion_claims | Truthfulness (1) | Lower is better | Truthfulness (0.0192%) | 0.00288% | 5 |
| Claude Opus 5.5 card — alignment and honestyfraud | Misuse resistance (1) | Lower is better | Misuse resistance (0.0106%) | 0.00106% | 5 |
| Claude Opus 5.5 card — alignment and honestyharmful_sexual_content | Misuse resistance (1) | Lower is better | Misuse resistance (0.0106%) | 0.00106% | 5 |
| Claude Opus 5.5 card — alignment and honestyharmful_system_prompt_compliance | Responsible agency (1) | Lower is better | Responsible agency (0.0203%) | 0.00305% | 5 |
| Claude Opus 5.5 card — alignment and honestyignoring_explicit_constraints | Responsible agency (1) | Lower is better | Responsible agency (0.0203%) | 0.00305% | 5 |
| Claude Opus 5.5 card — alignment and honestyimportant_omissions | Truthfulness (1) | Lower is better | Truthfulness (0.0192%) | 0.00288% | 5 |
| Claude Opus 5.5 card — alignment and honestyinput_hallucination | Truthfulness (1) | Lower is better | Truthfulness (0.0192%) | 0.00288% | 5 |
| Claude Opus 5.5 card — alignment and honestyinternalized_propaganda | Human rights (1) | Lower is better | Human rights (0.0411%) | 0.00617% | 5 |
| Claude Opus 5.5 card — alignment and honestyleaked_answer_undisclosed_use_rate_pct | Truthfulness (1) | Lower is better | Truthfulness (0.0288%) | 0.00432% | 5 |
| Claude Opus 5.5 card — alignment and honestymask_honesty_rate_pct | Truthfulness (1) | Higher is better | Truthfulness (0.0288%) | 0.00432% | 5 |
| Claude Opus 5.5 card — alignment and honestymilitary_grade_weapons | Misuse resistance (1) | Lower is better | Misuse resistance (0.0106%) | 0.00106% | 5 |
| Claude Opus 5.5 card — alignment and honestymotivated_reasoning | Truthfulness (1) | Lower is better | Truthfulness (0.0192%) | 0.00288% | 5 |
| Claude Opus 5.5 card — alignment and honestyoverrefusal | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.0733%) | 0.00733% | 5 |
| Claude Opus 5.5 card — alignment and honestypackage_registry_unencouraged_real_publish_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.0157%) | 0.00236% | 3 |
| Claude Opus 5.5 card — alignment and honestyradiological_nuclear_weapons_uplift | Misuse resistance (1) | Lower is better | Misuse resistance (0.0106%) | 0.00106% | 5 |
| Claude Opus 5.5 card — alignment and honestyreckless_tool_use | Responsible agency (1) | Lower is better | Responsible agency (0.0254%) | 0.00381% | 5 |
| Claude Opus 5.5 card — alignment and honestysandbox_moderate_or_higher_attempt_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.0182%) | 0.00273% | 4 |
| Claude Opus 5.5 card — alignment and honestysecurity_shortcut_for_expedience | Responsible agency (1) | Lower is better | Responsible agency (0.0254%) | 0.00381% | 5 |
| Claude Opus 5.5 card — alignment and honestyself_preservation | Responsible agency (1) | Lower is better | Responsible agency (0.0203%) | 0.00305% | 5 |
| Claude Opus 5.5 card — alignment and honestyself_serving_bias | Responsible agency (1) | Lower is better | Responsible agency (0.0203%) | 0.00305% | 5 |
| Claude Opus 5.5 card — alignment and honestysupporting_user_autonomy | Responsible agency (1) | Higher is better | Responsible agency (0.0203%) | 0.00305% | 5 |
| Claude Opus 5.5 card — alignment and honestysycophancy | Truthfulness (1) | Lower is better | Truthfulness (0.0192%) | 0.00288% | 5 |
| Claude Opus 5.5 card — alignment and honestyundermining_liberal_democracy | Human rights (1) | Lower is better | Human rights (0.0411%) | 0.00617% | 5 |
| Claude Opus 5.5 card — alignment and honestyunfaithful_thinking | Truthfulness (1) | Lower is better | Truthfulness (0.0192%) | 0.00288% | 5 |
| Claude Opus 5.5 card — alignment and honestyunprompted_boundary_probing | Responsible agency (1) | Lower is better | Responsible agency (0.0203%) | 0.00305% | 5 |
| Claude Opus 5.5 card — alignment and honestyunsanctioned_sandbox_escape_attempts | Responsible agency (1) | Lower is better | Responsible agency (0.0254%) | 0.00381% | 5 |
| Claude Opus 5.5 card — alignment and honestyunsanctioned_third_party_contact | Responsible agency (1) | Lower is better | Responsible agency (0.0203%) | 0.00305% | 5 |
| Claude Opus 5.5 card — alignment and honestyuser_deception | Truthfulness (1) | Lower is better | Truthfulness (0.0192%) | 0.00288% | 5 |
| Claude Opus 5.5 card — safety and refusal calibrationchild_api_multiturn_appropriate_rate_pct | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.0188%), Misuse resistance (0.0113%) | 0.00395% | 4 |
| Claude Opus 5.5 card — safety and refusal calibrationchild_claude_ai_multiturn_appropriate_rate_pct | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.0188%), Misuse resistance (0.0113%) | 0.00395% | 4 |
| Claude Opus 5.5 card — safety and refusal calibrationclaude_code_dual_use_benign_success_rate_pct | Benign helpfulness (1) | Higher is better | Benign helpfulness (0.0893%) | 0.00893% | 4 |
| Claude Opus 5.5 card — safety and refusal calibrationclaude_code_malicious_refusal_rate_pct | Misuse resistance (0.7), Responsible agency (0.3) | Higher is better | Misuse resistance (0.0135%), Responsible agency (0.0111%) | 0.00302% | 4 |
| Claude Opus 5.5 card — safety and refusal calibrationelection_api_harmless_rate_pct | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.0439%), Misuse resistance (0.0048%) | 0.00707% | 4 |
| Claude Opus 5.5 card — safety and refusal calibrationelection_api_multiturn_appropriate_rate_pct | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.0439%), Misuse resistance (0.0048%) | 0.00707% | 4 |
| Claude Opus 5.5 card — safety and refusal calibrationelection_api_overrefusal_rate_pct | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.0893%) | 0.00893% | 4 |
| Claude Opus 5.5 card — safety and refusal calibrationelection_claude_ai_harmless_rate_pct | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.0439%), Misuse resistance (0.0048%) | 0.00707% | 4 |
| Claude Opus 5.5 card — safety and refusal calibrationelection_claude_ai_multiturn_appropriate_rate_pct | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.0439%), Misuse resistance (0.0048%) | 0.00707% | 4 |
| Claude Opus 5.5 card — safety and refusal calibrationelection_claude_ai_overrefusal_rate_pct | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.0893%) | 0.00893% | 4 |
| Claude Opus 5.5 card — safety and refusal calibrationmalicious_computer_use_refusal_rate_pct | Misuse resistance (0.7), Responsible agency (0.3) | Higher is better | Misuse resistance (0.0135%), Responsible agency (0.0111%) | 0.00302% | 4 |
| Claude Opus 5.5 card — safety and refusal calibrationmultiturn_biological_weapons_api_appropriate_rate_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.0139%) | 0.00139% | 3 |
| Claude Opus 5.5 card — safety and refusal calibrationmultiturn_biological_weapons_claude_ai_appropriate_rate_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.0139%) | 0.00139% | 3 |
| Claude Opus 5.5 card — safety and refusal calibrationmultiturn_cyberattacks_api_appropriate_rate_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.0139%) | 0.00139% | 3 |
| Claude Opus 5.5 card — safety and refusal calibrationmultiturn_cyberattacks_claude_ai_appropriate_rate_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.0139%) | 0.00139% | 3 |
| Claude Opus 5.5 card — safety and refusal calibrationmultiturn_deadly_weapons_api_appropriate_rate_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.0139%) | 0.00139% | 3 |
| Claude Opus 5.5 card — safety and refusal calibrationmultiturn_deadly_weapons_claude_ai_appropriate_rate_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.0139%) | 0.00139% | 3 |
| Claude Opus 5.5 card — safety and refusal calibrationmultiturn_hate_and_discrimination_api_appropriate_rate_pct | Human rights (0.3), Fairness (0.5), Misuse resistance (0.2) | Higher is better | Human rights (0.0163%), Fairness (0.0341%), Misuse resistance (0.0028%) | 0.00613% | 3 |
| Claude Opus 5.5 card — safety and refusal calibrationmultiturn_hate_and_discrimination_claude_ai_appropriate_rate_pct | Human rights (0.3), Fairness (0.5), Misuse resistance (0.2) | Higher is better | Human rights (0.0163%), Fairness (0.0341%), Misuse resistance (0.0028%) | 0.00613% | 3 |
| Claude Opus 5.5 card — safety and refusal calibrationmultiturn_influence_operations_api_appropriate_rate_pct | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.038%), Misuse resistance (0.0042%) | 0.00612% | 3 |
| Claude Opus 5.5 card — safety and refusal calibrationmultiturn_influence_operations_claude_ai_appropriate_rate_pct | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.038%), Misuse resistance (0.0042%) | 0.00612% | 3 |
| Claude Opus 5.5 card — safety and refusal calibrationmultiturn_romance_scams_api_appropriate_rate_pct | Human rights (0.5), Misuse resistance (0.5) | Higher is better | Human rights (0.0272%), Misuse resistance (0.007%) | 0.00477% | 3 |
| Claude Opus 5.5 card — safety and refusal calibrationmultiturn_romance_scams_claude_ai_appropriate_rate_pct | Human rights (0.5), Misuse resistance (0.5) | Higher is better | Human rights (0.0272%), Misuse resistance (0.007%) | 0.00477% | 3 |
| Claude Opus 5.5 card — safety and refusal calibrationmultiturn_tracking_and_surveillance_api_appropriate_rate_pct | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.038%), Misuse resistance (0.0042%) | 0.00612% | 3 |
| Claude Opus 5.5 card — safety and refusal calibrationmultiturn_tracking_and_surveillance_claude_ai_appropriate_rate_pct | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.038%), Misuse resistance (0.0042%) | 0.00612% | 3 |
| Claude Opus 5.5 card — safety and refusal calibrationmultiturn_violent_extremism_api_appropriate_rate_pct | Human rights (0.5), Misuse resistance (0.5) | Higher is better | Human rights (0.0272%), Misuse resistance (0.007%) | 0.00477% | 3 |
| Claude Opus 5.5 card — safety and refusal calibrationmultiturn_violent_extremism_claude_ai_appropriate_rate_pct | Human rights (0.5), Misuse resistance (0.5) | Higher is better | Human rights (0.0272%), Misuse resistance (0.007%) | 0.00477% | 3 |
| Claude Opus 5.5 card — safety and refusal calibrationoverall_benign_api_overrefusal_rate_pct | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.0893%) | 0.00893% | 4 |
| Claude Opus 5.5 card — safety and refusal calibrationoverall_benign_claude_ai_overrefusal_rate_pct | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.0893%) | 0.00893% | 4 |
| Claude Opus 5.5 card — safety and refusal calibrationoverall_harmful_api_harmless_rate_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.0161%) | 0.00161% | 4 |
| Claude Opus 5.5 card — safety and refusal calibrationoverall_harmful_claude_ai_harmless_rate_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.0161%) | 0.00161% | 4 |
| Claude Opus 5.5 card — safety and refusal calibrationselfharm_api_multiturn_appropriate_rate_pct | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.0188%), Misuse resistance (0.0113%) | 0.00395% | 4 |
| Claude Opus 5.5 card — safety and refusal calibrationselfharm_claude_ai_multiturn_appropriate_rate_pct | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.0188%), Misuse resistance (0.0113%) | 0.00395% | 4 |
| Claude Sonnet 4.6 Overrefusalhigher_difficulty_overrefusal_rate | Benign helpfulness (1) | Lower is better | Benign helpfulness (1.05%) | 0.105% | 5 |
| Claude Sonnet 4.6 Overrefusaloverall_overrefusal_rate | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.875%) | 0.0875% | 5 |
| Claude Sonnet 4.6 User Wellbeingchild_benign_refusal_rate | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.783%) | 0.0783% | 4 |
| Claude Sonnet 4.6 User Wellbeingchild_multiturn_appropriate_rate | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.369%), Misuse resistance (0.0406%) | 0.0595% | 4 |
| Claude Sonnet 4.6 User Wellbeingchild_violative_harmless_rate | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.369%), Misuse resistance (0.0406%) | 0.0595% | 4 |
| Claude Sonnet 4.6 User Wellbeingselfharm_benign_refusal_rate | Benign helpfulness (1) | Lower is better | Benign helpfulness (0.626%) | 0.0626% | 4 |
| Claude Sonnet 4.6 User Wellbeingselfharm_harmless_rate | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.132%), Misuse resistance (0.079%) | 0.0277% | 4 |
| Claude Sonnet 4.6 User Wellbeingselfharm_multiturn_appropriate_rate | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.158%), Misuse resistance (0.0947%) | 0.0332% | 4 |
| Claude system cards — Gray Swan Q1+Q2 indirect prompt injection k=15attack_success_probability_k15_pct | Responsible agency (1) | Lower is better | Responsible agency (1.11%) | 0.167% | 12 |
| CMoralEvalfamilial_morality | Misuse resistance (1) | Higher is better | Misuse resistance (0.247%) | 0.0247% | 26 |
| CMoralEvalinternet_ethics | Misuse resistance (1) | Higher is better | Misuse resistance (0.247%) | 0.0247% | 26 |
| CMoralEvalpersonal_morality | Misuse resistance (1) | Higher is better | Misuse resistance (0.247%) | 0.0247% | 26 |
| CMoralEvalprofessional_ethics | Misuse resistance (1) | Higher is better | Misuse resistance (0.247%) | 0.0247% | 26 |
| CMoralEvalsocial_morality | Misuse resistance (1) | Higher is better | Misuse resistance (0.247%) | 0.0247% | 26 |
| COMPL-AI AI-Identity Disclosurescore | Truthfulness (1) | Higher is better | Truthfulness (0.55%) | 0.0824% | 14 |
| COMPL-AI LLM RuLES Multi-Turn Rule Followingscore | Misuse resistance (1) | Higher is better | Misuse resistance (0.302%) | 0.0302% | 14 |
| COMPL-AI TensorTrust Goal-Hijacking Resistancescore | Responsible agency (1) | Higher is better | Responsible agency (0.561%) | 0.0841% | 13 |
| Concordia — Agentic-Misalignmentsafety_score | Responsible agency (1) | Higher is better | Responsible agency (0.686%) | 0.103% | 54 |
| Concordia — AirBench-Deceptionsafety_score | Misuse resistance (1) | Higher is better | Misuse resistance (0.385%) | 0.0385% | 63 |
| Concordia — AirBench-Manipulationsafety_score | Human rights (1) | Higher is better | Human rights (1.41%) | 0.212% | 56 |
| Concordia — AirBench-SecurityRiskssafety_score | Misuse resistance (1) | Higher is better | Misuse resistance (0.385%) | 0.0385% | 63 |
| Concordia — APEsafety_score | Human rights (0.5), Responsible agency (0.5) | Higher is better | Human rights (0.701%), Responsible agency (0.346%) | 0.157% | 55 |
| Concordia — CyberSecEval2-PromptInjectionsafety_score | Responsible agency (1) | Higher is better | Responsible agency (0.741%) | 0.111% | 63 |
| Concordia — DarkBenchsafety_score | Human rights (0.5), Truthfulness (0.5) | Higher is better | Human rights (0.701%), Truthfulness (0.327%) | 0.154% | 55 |
| Concordia — Fortress-Biologicalsafety_score | Misuse resistance (1) | Higher is better | Misuse resistance (0.356%) | 0.0356% | 54 |
| Concordia — Fortress-Chemicalsafety_score | Misuse resistance (1) | Higher is better | Misuse resistance (0.356%) | 0.0356% | 54 |
| Concordia — Fortress-Privacy/Scamssafety_score | Misuse resistance (1) | Higher is better | Misuse resistance (0.356%) | 0.0356% | 54 |
| Concordia — FRT-AirBench-Manipulationsafety_score | Human rights (1) | Higher is better | Human rights (1.27%) | 0.19% | 45 |
| Concordia — FRT-AirBench-SecurityRiskssafety_score | Misuse resistance (1) | Higher is better | Misuse resistance (0.325%) | 0.0325% | 45 |
| Concordia — FRT-SciKnowEval-BiologicalHarmfulQAsafety_score | Misuse resistance (1) | Higher is better | Misuse resistance (0.325%) | 0.0325% | 45 |
| Concordia — FRT-SOSBench-Chemsafety_score | Misuse resistance (1) | Higher is better | Misuse resistance (0.325%) | 0.0325% | 45 |
| Concordia — MASKsafety_score | Truthfulness (1) | Higher is better | Truthfulness (0.694%) | 0.104% | 62 |
| Concordia — SciKnowEval-BiologicalHarmfulQAsafety_score | Misuse resistance (1) | Higher is better | Misuse resistance (0.385%) | 0.0385% | 63 |
| Concordia — SciKnowEval-ChemicalHarmfulQAsafety_score | Misuse resistance (1) | Higher is better | Misuse resistance (0.385%) | 0.0385% | 63 |
| Concordia — Shutdown-Resistancesafety_score | Responsible agency (1) | Higher is better | Responsible agency (0.679%) | 0.102% | 53 |
| Concordia — SOSBench-Biosafety_score | Misuse resistance (1) | Higher is better | Misuse resistance (0.385%) | 0.0385% | 63 |
| Concordia — SOSBench-Chemsafety_score | Misuse resistance (1) | Higher is better | Misuse resistance (0.385%) | 0.0385% | 63 |
| Confabulationsconfabulation_rate | Truthfulness (1) | Lower is better | Truthfulness (4.77%) | 0.715% | 52 |
| Constitutional Following — Anthropic Constitutionconstitutional_following_score | Responsible agency (1) | Higher is better | Responsible agency (0.617%) | 0.0926% | 7 |
| Constitutional Following — OpenAI Model Specconstitutional_following_score | Responsible agency (1) | Higher is better | Responsible agency (0.617%) | 0.0926% | 7 |
| Contextual MoralChoicehuman_agreement | Misuse resistance (1) | Higher is better | Misuse resistance (1.14%) | 0.114% | 22 |
| CRiskEvaldeception_willingness | Truthfulness (1) | Lower is better | Truthfulness (0.26%) | 0.0389% | 17 |
| CRiskEvaldesire_for_resource | Responsible agency (1) | Lower is better | Responsible agency (0.275%) | 0.0412% | 17 |
| CRiskEvalharmful_goal | Misuse resistance (1) | Lower is better | Misuse resistance (0.143%) | 0.0143% | 17 |
| CRiskEvalimprovement_intent | Responsible agency (1) | Lower is better | Responsible agency (0.275%) | 0.0412% | 17 |
| CRiskEvalmalicious_coordination | Responsible agency (1) | Lower is better | Responsible agency (0.275%) | 0.0412% | 17 |
| CRiskEvalself_preservation | Responsible agency (1) | Lower is better | Responsible agency (0.275%) | 0.0412% | 17 |
| CRiskEvalsituational_awareness | Responsible agency (1) | Lower is better | Responsible agency (0.275%) | 0.0412% | 17 |
| CValuesscore | Misuse resistance (1) | Higher is better | Misuse resistance (1.15%) | 0.115% | 10 |
| DecodingTrustmachine_ethics | Misuse resistance (1) | Higher is better | Misuse resistance (0.114%) | 0.0114% | 8 |
| DecodingTruststereotype_bias | Fairness (1) | Higher is better | Fairness (1.4%) | 0.14% | 8 |
| DecodingTrusttoxicity | Misuse resistance (1) | Higher is better | Misuse resistance (0.286%) | 0.0286% | 8 |
| DelusionEvaldelusional_prevalence_pct | Truthfulness (1) | Lower is better | Truthfulness (0.529%) | 0.0793% | 16 |
| DelusionEvaldiscourages_harm_prevalence_pct | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.34%), Misuse resistance (0.204%) | 0.0714% | 16 |
| DelusionEvalfacilitates_harm_prevalence_pct | Human rights (0.3), Misuse resistance (0.7) | Lower is better | Human rights (0.34%), Misuse resistance (0.204%) | 0.0714% | 16 |
| DelusionEvalrelationship_prevalence_pct | Human rights (0.7), Responsible agency (0.3) | Lower is better | Human rights (0.794%), Responsible agency (0.168%) | 0.144% | 16 |
| DelusionEvalsycophancy_prevalence_pct | Truthfulness (1) | Lower is better | Truthfulness (0.529%) | 0.0793% | 16 |
| Do-Not-Answerhuman_harmlessness_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.557%) | 0.0557% | 6 |
| DSPSafeBenchscore | Misuse resistance (1) | Higher is better | Misuse resistance (1.68%) | 0.168% | 12 |
| DystopiaBenchbasaglia_score | Human rights (1) | Lower is better | Human rights (2.67%) | 0.401% | 50 |
| DystopiaBenchbaudrillard_score | Human rights (0.33), Truthfulness (0.67) | Lower is better | Human rights (0.445%), Truthfulness (0.416%) | 0.129% | 50 |
| DystopiaBenchhuxley_score | Human rights (1) | Lower is better | Human rights (2.67%) | 0.401% | 50 |
| DystopiaBenchlaguardia_score | Human rights (1) | Lower is better | Human rights (2.67%) | 0.401% | 50 |
| DystopiaBenchorwell_score | Human rights (0.67), Truthfulness (0.33) | Lower is better | Human rights (0.892%), Truthfulness (0.207%) | 0.165% | 50 |
| DystopiaBenchpetrov_score | Human rights (1) | Lower is better | Human rights (2.67%) | 0.401% | 50 |
| Emergent Collusionhigh_illegality_game_rate | Responsible agency (1) | Lower is better | Responsible agency (3.36%) | 0.505% | 13 |
| Enkrypt AI Safety Leaderboardbias_attack_non_success_rate | Human rights (0.8), Misuse resistance (0.2) | Higher is better | Human rights (6.21%), Misuse resistance (0.398%) | 0.972% | 270 |
| Enkrypt AI Safety Leaderboardcbrn_attack_non_success_rate | Misuse resistance (0.4), Responsible agency (0.6) | Higher is better | Misuse resistance (0.797%), Responsible agency (2.3%) | 0.425% | 270 |
| Enkrypt AI Safety Leaderboardharmful_attack_non_success_rate | Human rights (0.15), Misuse resistance (0.7), Responsible agency (0.15) | Higher is better | Human rights (1.16%), Misuse resistance (1.39%), Responsible agency (0.575%) | 0.4% | 270 |
| Enkrypt AI Safety Leaderboardtoxicity_attack_non_success_rate | Human rights (0.35), Misuse resistance (0.65) | Higher is better | Human rights (2.71%), Misuse resistance (1.29%) | 0.535% | 268 |
| Every Model Cheats — Cybench Cheat Propensitybaseline_cheat_propensity_pct | Responsible agency (1) | Lower is better | Responsible agency (1.09%) | 0.164% | 22 |
| Every Model Cheats — Cybench Cheat Propensitysevere_anti_cheat_cheat_propensity_pct | Responsible agency (1) | Lower is better | Responsible agency (1.09%) | 0.164% | 22 |
| Every Model Cheats — Cybench Cheat Propensitystandard_anti_cheat_cheat_propensity_pct | Responsible agency (1) | Lower is better | Responsible agency (1.09%) | 0.164% | 22 |
| Fake Alignment (FINE)multiple_choice_safe_decision_rate | Human rights (0.4), Fairness (0.2), Misuse resistance (0.4) | Higher is better | Human rights (0.849%), Fairness (0.532%), Misuse resistance (0.218%) | 0.202% | 14 |
| Fake Alignment (FINE)open_ended_safe_response_rate | Human rights (0.4), Fairness (0.2), Misuse resistance (0.4) | Higher is better | Human rights (1.27%), Fairness (0.798%), Misuse resistance (0.327%) | 0.303% | 14 |
| FinEval 6.0 Safety Awarenesssafety_awareness_score | Misuse resistance (1) | Higher is better | Misuse resistance (0.343%) | 0.0343% | 8 |
| FinEval Financial Security Knowledgefinancial_security_accuracy_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.528%) | 0.0529% | 19 |
| FlagEval Safety and Valuesa1_qualified_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.309%) | 0.0309% | 18 |
| FlagEval Safety and Valuesa2_qualified_rate | Fairness (1) | Higher is better | Fairness (1.51%) | 0.151% | 18 |
| FlagEval Safety and Valuesa3_qualified_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.309%) | 0.0309% | 18 |
| FlagEval Safety and Valuesa4_qualified_rate | Human rights (1) | Higher is better | Human rights (1.2%) | 0.18% | 18 |
| FlagEval Safety and Valuesa5_qualified_rate | Benign helpfulness (1) | Higher is better | Benign helpfulness (2.14%) | 0.214% | 18 |
| FLAMESdata_protection | Human rights (1) | Higher is better | Human rights (1.36%) | 0.204% | 13 |
| FLAMESfairness | Fairness (1) | Higher is better | Fairness (1.71%) | 0.171% | 13 |
| FLAMESlegality | Misuse resistance (1) | Higher is better | Misuse resistance (0.35%) | 0.035% | 13 |
| FLAMESmorality | Misuse resistance (1) | Higher is better | Misuse resistance (0.35%) | 0.035% | 13 |
| FLAMESsafety | Misuse resistance (1) | Higher is better | Misuse resistance (0.35%) | 0.035% | 13 |
| FORTRESSaverage_risk_score | Misuse resistance (1) | Lower is better | Misuse resistance (1.88%) | 0.188% | 60 |
| FORTRESSover_refusal_score | Benign helpfulness (1) | Lower is better | Benign helpfulness (12.9%) | 1.29% | 59 |
| Google Gemini 2.5 Flash Model Cardtext_safety_delta | Misuse resistance (1) | Lower is better | Misuse resistance (0.514%) | 0.0514% | 2 |
| Google Gemini 2.5 Flash-Lite Model Cardtext_safety_delta | Misuse resistance (1) | Lower is better | Misuse resistance (0.514%) | 0.0514% | 2 |
| Google Gemini 3.8 launch — Gray Swan indirect prompt injection k=15attack_success_probability_k15_pct | Responsible agency (1) | Lower is better | Responsible agency (3.61%) | 0.542% | 15 |
| Governance Decay under Passive Context Compactiongovernance_retention_score | Responsible agency (1) | Higher is better | Responsible agency (1.23%) | 0.185% | 7 |
| GPT 6 Astra system-card alignment evaluationsattempted_retry_or_circumvention_after_denial_pct | Responsible agency (1) | Lower is better | Responsible agency (0.0606%) | 0.0091% | 2 |
| GPT 6 Astra system-card alignment evaluationsexploitgym_honeypot_attack_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.0728%) | 0.0109% | 2 |
| GPT 6 Astra system-card alignment evaluationsfalsified_data_labels_pct | Truthfulness (1) | Lower is better | Truthfulness (0.0344%) | 0.00515% | 2 |
| GPT 6 Astra system-card alignment evaluationsfollowed_unauthorized_agent_instructions_pct | Responsible agency (1) | Lower is better | Responsible agency (0.0858%) | 0.0129% | 4 |
| GPT 6 Astra system-card alignment evaluationsinternal_circumvention_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.0606%) | 0.0091% | 2 |
| GPT 6 Astra system-card alignment evaluationsinternal_computer_use_safety_autoreview_error_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.0606%) | 0.0091% | 2 |
| GPT 6 Astra system-card alignment evaluationsinternal_computer_use_safety_error_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.115%) | 0.0173% | 5 |
| GPT 6 Astra system-card alignment evaluationsinternal_hallucination_rate_pct | Truthfulness (1) | Lower is better | Truthfulness (0.0458%) | 0.00687% | 2 |
| GPT 6 Astra system-card alignment evaluationsoverall_misaligned_outcome_base_pct | Responsible agency (1) | Lower is better | Responsible agency (0.103%) | 0.0154% | 4 |
| GPT 6 Astra system-card alignment evaluationsoverall_misaligned_outcome_confirmation_pct | Responsible agency (1) | Lower is better | Responsible agency (0.0858%) | 0.0129% | 4 |
| GPT 6 Astra system-card alignment evaluationsseverity_1_or_2_misalignment_flags_pct | Responsible agency (1) | Lower is better | Responsible agency (0.0728%) | 0.0109% | 2 |
| GPT 6 Astra system-card alignment evaluationsseverity_3_plus_misalignment_flags_pct | Responsible agency (1) | Lower is better | Responsible agency (0.097%) | 0.0146% | 2 |
| GPT 6 Astra system-card alignment evaluationsunwanted_persistence_after_warning_pct | Responsible agency (1) | Lower is better | Responsible agency (0.0606%) | 0.0091% | 2 |
| GPT-5.6 system cardconnectors_injection_resistance | Responsible agency (1) | Higher is better | Responsible agency (0.198%) | 0.0298% | 7 |
| GPT-5.6 system cardemotional_reliance | Human rights (0.7), Responsible agency (0.3) | Higher is better | Human rights (0.438%), Responsible agency (0.0926%) | 0.0795% | 7 |
| GPT-5.6 system cardextremism_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.0859%) | 0.00859% | 7 |
| GPT-5.6 system cardgore_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.0859%) | 0.00859% | 7 |
| GPT-5.6 system cardharm_overall_pct | Fairness (1) | Lower is better | Fairness (0.42%) | 0.042% | 7 |
| GPT-5.6 system cardhate_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.0859%) | 0.00859% | 7 |
| GPT-5.6 system cardmental_health | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.365%), Misuse resistance (0.0401%) | 0.0587% | 7 |
| GPT-5.6 system cardnonviolent_illicit_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.0859%) | 0.00859% | 7 |
| GPT-5.6 system cardsearch_function_calling_injection_resistance | Responsible agency (1) | Higher is better | Responsible agency (0.184%) | 0.0276% | 6 |
| GPT-5.6 system cardself_harm | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.156%), Misuse resistance (0.0936%) | 0.0328% | 7 |
| GPT-5.6 system cardself_harm_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.0859%) | 0.00859% | 7 |
| GPT-5.6 system cardsexual_minors_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.0859%) | 0.00859% | 7 |
| GPT-5.6 system cardsexual_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.0687%) | 0.00687% | 7 |
| GPT-5.6 system cardviolent_illicit_not_unsafe | Misuse resistance (1) | Higher is better | Misuse resistance (0.0859%) | 0.00859% | 7 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsagentic_chat_plugins_safe_rate | Responsible agency (1) | Higher is better | Responsible agency (0.019%) | 0.00285% | 4 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsagentic_codex_age_restricted_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.0099%) | 0.000986% | 4 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsagentic_codex_nonviolent_wrongdoing_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.0099%) | 0.000986% | 4 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsagentic_codex_personal_data_safe_rate | Responsible agency (1) | Higher is better | Responsible agency (0.019%) | 0.00285% | 4 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsagentic_codex_self_harm_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.0099%) | 0.000986% | 4 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsagentic_codex_violent_wrongdoing_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.0099%) | 0.000986% | 4 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsagentic_redteam_chat_plugins_safe_rate | Responsible agency (1) | Higher is better | Responsible agency (0.019%) | 0.00285% | 4 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsagentic_redteam_codex_safe_rate | Responsible agency (1) | Higher is better | Responsible agency (0.019%) | 0.00285% | 4 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsattempted_retry_or_circumvention_after_denial_pct_2026_09_22 | Responsible agency (1) | Lower is better | Responsible agency (0.0212%) | 0.00318% | 5 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsbiology_benign_not_overrefuse_rate | Benign helpfulness (1) | Higher is better | Benign helpfulness (0.0838%) | 0.00838% | 6 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsbiology_dual_use_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.0121%) | 0.00121% | 6 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsbiology_severe_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.0121%) | 0.00121% | 6 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsbroken_search_unacknowledged_limitation_pct_2026_09_22 | Truthfulness (1) | Lower is better | Truthfulness (0.02%) | 0.00301% | 5 |
| GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_extremism_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.0065%) | 0.000652% | 7 |
| GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_gore_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.0065%) | 0.000652% | 7 |
| GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_hate_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.0065%) | 0.000652% | 7 |
| GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_nonviolent_illicit_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.0065%) | 0.000652% | 7 |
| GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_self_harm_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.0065%) | 0.000652% | 7 |
| GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_sexual_minors_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.0065%) | 0.000652% | 7 |
| GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_sexual_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.0065%) | 0.000652% | 7 |
| GPT-6 Sol/Luna system card — safety and updated alignment testschallenging_violent_illicit_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.0065%) | 0.000652% | 7 |
| GPT-6 Sol/Luna system card — safety and updated alignment testscoding_deception_misrepresentation_pct_2026_09_22 | Truthfulness (1) | Lower is better | Truthfulness (0.02%) | 0.00301% | 5 |
| GPT-6 Sol/Luna system card — safety and updated alignment testscyber_production_chat_safe_score | Misuse resistance (1) | Higher is better | Misuse resistance (0.0121%) | 0.00121% | 6 |
| GPT-6 Sol/Luna system card — safety and updated alignment testscyber_semisynthetic_agentic_safe_score | Misuse resistance (1) | Higher is better | Misuse resistance (0.011%) | 0.0011% | 5 |
| GPT-6 Sol/Luna system card — safety and updated alignment testscyber_synthetic_agentic_safe_score | Misuse resistance (1) | Higher is better | Misuse resistance (0.0121%) | 0.00121% | 6 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsdynamic_emotional_reliance_safe_rate | Responsible agency (1) | Higher is better | Responsible agency (0.0126%) | 0.00188% | 7 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsdynamic_mental_health_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.0065%) | 0.000652% | 7 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsdynamic_self_harm_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.0065%) | 0.000652% | 7 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsinstruction_hierarchy_defender_success_pct | Responsible agency (1) | Higher is better | Responsible agency (0.0082%) | 0.00123% | 3 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsseverity_1_or_2_misalignment_flags_pct_2026_09_22 | Responsible agency (1) | Lower is better | Responsible agency (0.0134%) | 0.00201% | 2 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsseverity_3_plus_misalignment_flags_pct_2026_09_22 | Responsible agency (1) | Lower is better | Responsible agency (0.0134%) | 0.00201% | 2 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsstatic_jailbreak_bio_high_risk_defender_success_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.011%) | 0.0011% | 5 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsstatic_jailbreak_bio_severe_defender_success_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.011%) | 0.0011% | 5 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsstatic_jailbreak_cyber_defender_success_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.011%) | 0.0011% | 5 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsstatic_jailbreak_violence_moderate_defender_success_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.011%) | 0.0011% | 5 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsstatic_jailbreak_violence_severe_defender_success_pct | Misuse resistance (1) | Higher is better | Misuse resistance (0.011%) | 0.0011% | 5 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsu18_age_restricted_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.0065%) | 0.000652% | 7 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsu18_eating_disorders_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.0065%) | 0.000652% | 7 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsu18_emotional_reliance_safe_rate | Responsible agency (1) | Higher is better | Responsible agency (0.0126%) | 0.00188% | 7 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsu18_gore_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.0065%) | 0.000652% | 7 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsu18_self_harm_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.0065%) | 0.000652% | 7 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsu18_sexual_content_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.0065%) | 0.000652% | 7 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsunwanted_persistence_after_warning_pct_2026_09_22 | Responsible agency (1) | Lower is better | Responsible agency (0.0212%) | 0.00318% | 5 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsvision_erotic_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.006%) | 0.000604% | 6 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsvision_extremism_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.006%) | 0.000604% | 6 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsvision_hate_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.006%) | 0.000604% | 6 |
| GPT-6 Sol/Luna system card — safety and updated alignment testsvision_self_harm_safe_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.006%) | 0.000604% | 6 |
| Gray Swan indirect prompt injection (15 attempts)attack_success_probability_k15_pct | Responsible agency (1) | Lower is better | Responsible agency (2.88%) | 0.433% | 13 |
| HalluVerse-M3 Hallucination Recognitionhallucination_recognition_accuracy | Truthfulness (1) | Higher is better | Truthfulness (1.65%) | 0.247% | 14 |
| HarmBenchdr | Misuse resistance (1) | Lower is better | Misuse resistance (0.535%) | 0.0535% | 28 |
| HarmVideoBenchharmful_video_safety_recognition_reasoning | Misuse resistance (1) | Higher is better | Misuse resistance (1.06%) | 0.106% | 19 |
| HELM Classic RealToxicityPromptsnon_toxic_prompt_toxic_fraction | Misuse resistance (1) | Lower is better | Misuse resistance (0.767%) | 0.0767% | 40 |
| HELM Classic RealToxicityPromptstoxic_prompt_toxic_fraction | Misuse resistance (1) | Lower is better | Misuse resistance (0.767%) | 0.0767% | 40 |
| HELM Safetyanthropic_red_team | Misuse resistance (1) | Higher is better | Misuse resistance (4.34%) | 0.434% | 80 |
| HELM Safetybbq | Fairness (1) | Higher is better | Fairness (15.9%) | 1.59% | 80 |
| HELM Safetyharmbench | Misuse resistance (1) | Higher is better | Misuse resistance (0.904%) | 0.0904% | 80 |
| HELM Safetysimple_safety_tests | Misuse resistance (1) | Higher is better | Misuse resistance (4.34%) | 0.434% | 80 |
| HELM Safetyxstest | Benign helpfulness (0.56), Misuse resistance (0.44) | Higher is better | Benign helpfulness (4.18%), Misuse resistance (0.481%) | 0.467% | 80 |
| HUMAINE Trust, Ethics and Safetytrust_ethics_safety_score | Truthfulness (0.3), Misuse resistance (0.7) | Higher is better | Truthfulness (1.94%), Misuse resistance (2.49%) | 0.541% | 54 |
| Human Pathogen Capabilities Test (HPCT) — overall refusalrefusal_rate_pct | Misuse resistance (1) | Higher is better | Misuse resistance (2.01%) | 0.201% | 69 |
| Humanity's Last Exam RMS calibration error (Scale Labs)calibrationError | Truthfulness (1) | Lower is better | Truthfulness (1.43%) | 0.214% | 42 |
| HyperCLOVA X Toxic Continuation Panelskold_toxic_count | Misuse resistance (1) | Lower is better | Misuse resistance (0.241%) | 0.0241% | 7 |
| HyperCLOVA X Toxic Continuation Panelskold_toxicity | Misuse resistance (1) | Lower is better | Misuse resistance (0.301%) | 0.0301% | 7 |
| HyperCLOVA X Toxic Continuation Panelsrealt-toxicprompts_toxic_count | Misuse resistance (1) | Lower is better | Misuse resistance (0.18%) | 0.018% | 7 |
| HyperCLOVA X Toxic Continuation Panelsrealt-toxicprompts_toxicity | Misuse resistance (1) | Lower is better | Misuse resistance (0.241%) | 0.0241% | 7 |
| IndoBias-Pairs — parity-aware culturally grounded biasparity_score | Fairness (1) | Higher is better | Fairness (6.04%) | 0.604% | 26 |
| Inkling-Small model card — FORTRESSbenign_answer_rate | Benign helpfulness (1) | Higher is better | Benign helpfulness (1.77%) | 0.177% | 10 |
| Inkling-Small model card — FORTRESSharmful_refusal_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.256%) | 0.0256% | 10 |
| Inkling-Small model card — StrongREJECTsafety_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.256%) | 0.0256% | 10 |
| JailBenchjailbreak_success_rate | Misuse resistance (1) | Lower is better | Misuse resistance (0.907%) | 0.0907% | 14 |
| JuICE Cultural-Error Span Detectionf1 | Fairness (1) | Higher is better | Fairness (3.75%) | 0.375% | 10 |
| KIDBench Implicit Child Cueimplicit_child_cue_total_mean | Human rights (0.2), Benign helpfulness (0.4), Misuse resistance (0.4) | Higher is better | Human rights (1.02%), Benign helpfulness (3.64%), Misuse resistance (0.525%) | 0.57% | 13 |
| kindbench v0.1.0 psychological safety rankingemotional_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.383%) | 0.0383% | 10 |
| kindbench v0.1.0 psychological safety rankingidentity_collapse | Responsible agency (1) | Higher is better | Responsible agency (0.738%) | 0.111% | 10 |
| kindbench v0.1.0 psychological safety rankingsycophancy_spine | Truthfulness (1) | Higher is better | Truthfulness (0.697%) | 0.105% | 10 |
| kindbench v0.1.0 psychological safety rankingvalue_integrity | Fairness (1) | Higher is better | Fairness (1.87%) | 0.187% | 10 |
| Large-scale Moral Machine experiment on LLMshuman_choice_distance | Nonhuman welfare (0.11), Human rights (0.45), Fairness (0.44) | Lower is better | Nonhuman welfare (1.17%), Human rights (2.63%), Fairness (3.29%) | 1.02% | 39 |
| LiveSecBenchethics | Fairness (1) | Higher is better | Fairness (1.67%) | 0.167% | 43 |
| LiveSecBenchfactuality | Truthfulness (1) | Higher is better | Truthfulness (1.03%) | 0.155% | 43 |
| LiveSecBenchlegality | Misuse resistance (1) | Higher is better | Misuse resistance (0.568%) | 0.0568% | 43 |
| LiveSecBenchprivacy | Human rights (1) | Higher is better | Human rights (2.21%) | 0.332% | 43 |
| LiveSecBenchpsychological_health | Human rights (1) | Higher is better | Human rights (1.33%) | 0.199% | 43 |
| LLM Ethics Benchmarkscore | Human rights (1) | Higher is better | Human rights (2.11%) | 0.317% | 5 |
| M3-SafetyBenchoverall_score | Misuse resistance (1) | Higher is better | Misuse resistance (1.06%) | 0.106% | 19 |
| MACHIAVELLIdeception_relative_random_pct | Truthfulness (1) | Lower is better | Truthfulness (3.12%) | 0.467% | 50 |
| Manager Coercion Benchcoercion_ladder_depth | Responsible agency (1) | Lower is better | Responsible agency (3.13%) | 0.47% | 15+30 |
| Manager Coercion Benchfabrication_rate | Truthfulness (1) | Lower is better | Truthfulness (1.71%) | 0.256% | 15 |
| MANTAAWMS | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (5.87%) | 1.47% | 12 |
| MANTAAWVS | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (5.87%) | 1.47% | 12 |
| MASKlying_probability_pct | Truthfulness (1) | Lower is better | Truthfulness (3.21%) | 0.481% | 53 |
| MedSafetyBenchmedical_safety_score | Misuse resistance (1) | Higher is better | Misuse resistance (2.66%) | 0.266% | 30 |
| MentalHealthBenchoverall_model_performance_pct | Benign helpfulness (0.5), Misuse resistance (0.5) | Higher is better | Benign helpfulness (5.2%), Misuse resistance (0.75%) | 0.595% | 17 |
| Microsoft Phi Safety Panelsharmful_continuation | Misuse resistance (1) | Lower is better | Misuse resistance (0.274%) | 0.0274% | 10 |
| Microsoft Phi Safety Panelsharmful_summarization | Misuse resistance (1) | Lower is better | Misuse resistance (0.274%) | 0.0274% | 10 |
| Microsoft Phi Safety Panelsjailbreak | Misuse resistance (1) | Lower is better | Misuse resistance (0.329%) | 0.0329% | 10 |
| Microsoft Phi Safety Panelsthird_party_harm | Human rights (1) | Lower is better | Human rights (1.07%) | 0.16% | 10 |
| MonitoringBench Full-Trajectory Monitorfull_trajectory_catch_rate_at_1pct_fpr_percent | Responsible agency (1) | Higher is better | Responsible agency (1.68%) | 0.252% | 13 |
| MORUscore | Nonhuman welfare (0.72), Human rights (0.1), Responsible agency (0.18) | Higher is better | Nonhuman welfare (4.37%), Human rights (0.354%), Responsible agency (0.303%) | 1.19% | 13 |
| MT-JailBench CrescendoXsafety_score | Misuse resistance (1) | Higher is better | Misuse resistance (0.185%) | 0.0185% | 21 |
| MuPPET Contextual Privacymultiparty_contextual_privacy_score | Human rights (1) | Higher is better | Human rights (2.5%) | 0.375% | 7 |
| NESSiE Necessary Safety Benchmarksh_score | Benign helpfulness (0.5), Misuse resistance (0.5) | Higher is better | Benign helpfulness (6.73%), Misuse resistance (0.97%) | 0.77% | 16 |
| ODCV-Benchaverage_severity | Responsible agency (1) | Lower is better | Responsible agency (2.02%) | 0.303% | 12 |
| ODCV-Benchmisalignment_rate | Responsible agency (1) | Lower is better | Responsible agency (2.02%) | 0.303% | 12 |
| Olam Diplomacy — Broken Promise Ratebroken_promise_rate_pct | Truthfulness (1) | Lower is better | Truthfulness (1.92%) | 0.288% | 19 |
| Olam Social Poker — Social Lie Ratesocial_lie_rate_per_10000_turns | Truthfulness (1) | Lower is better | Truthfulness (2.16%) | 0.324% | 24 |
| Open LLM Safety Indexjailbreakbench_safety_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.556%) | 0.0556% | 21 |
| Open LLM Safety Indexstrongreject_safety_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.556%) | 0.0556% | 21 |
| OpenAgentSafetyllm_judge_safety_vulnerable | Responsible agency (1) | Lower is better | Responsible agency (0.926%) | 0.139% | 7 |
| OpenAgentSafetyrule_based_safety_vulnerable | Responsible agency (1) | Lower is better | Responsible agency (0.772%) | 0.116% | 7 |
| OpenAgentSafetysuccessful_completion | Responsible agency (1) | Higher is better | Responsible agency (0.772%) | 0.116% | 7 |
| OpenAI GPT-4o System Cardspeaker_privacy_refusal_accuracy | Human rights (1) | Higher is better | Human rights (2.01%) | 0.301% | 2 |
| OpenAI GPT-5 System Cardharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.514%) | 0.0514% | 2 |
| OpenAI GPT-5.3 Dynamic Wellbeingemotional_reliance | Human rights (0.7), Responsible agency (0.3) | Higher is better | Human rights (0.286%), Responsible agency (0.0606%) | 0.0521% | 3 |
| OpenAI GPT-5.3 Dynamic Wellbeingmental_health | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.239%), Misuse resistance (0.0263%) | 0.0384% | 3 |
| OpenAI GPT-5.3 Dynamic Wellbeingself_harm | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.102%), Misuse resistance (0.0613%) | 0.0215% | 3 |
| OpenAI GPT-5.4 Dynamic Wellbeingemotional_reliance | Human rights (0.7), Responsible agency (0.3) | Higher is better | Human rights (0.286%), Responsible agency (0.0606%) | 0.0521% | 3 |
| OpenAI GPT-5.4 Dynamic Wellbeingmental_health | Human rights (0.7), Misuse resistance (0.3) | Higher is better | Human rights (0.239%), Misuse resistance (0.0263%) | 0.0384% | 3 |
| OpenAI GPT-5.4 Dynamic Wellbeingself_harm | Human rights (0.3), Misuse resistance (0.7) | Higher is better | Human rights (0.102%), Misuse resistance (0.0613%) | 0.0215% | 3 |
| OpenAI GPT-5.4 First-Person Fairnessharm_overall | Fairness (1) | Lower is better | Fairness (0.604%) | 0.0604% | 3 |
| OpenAI GPT-5.4 Property Preservationdestructive_action_avoidance | Human rights (0.3), Responsible agency (0.7) | Higher is better | Human rights (0.173%), Responsible agency (0.2%) | 0.056% | 3 |
| OpenAI GPT-5.4 Property Preservationuser_work_preserved | Human rights (0.4), Responsible agency (0.6) | Higher is better | Human rights (0.231%), Responsible agency (0.171%) | 0.0603% | 3 |
| OpenAI GPT-5.4 User Confirmationsfinancial_transaction | Human rights (0.3), Responsible agency (0.7) | Higher is better | Human rights (0.173%), Responsible agency (0.2%) | 0.056% | 3 |
| OpenAI GPT-5.4 User Confirmationsgeneral_confirmation | Human rights (0.3), Responsible agency (0.7) | Higher is better | Human rights (0.145%), Responsible agency (0.166%) | 0.0466% | 3 |
| OpenAI GPT-5.4 User Confirmationshigh_stakes_communication | Human rights (0.3), Responsible agency (0.7) | Higher is better | Human rights (0.173%), Responsible agency (0.2%) | 0.056% | 3 |
| OpenAI o3 and o4-mini System Cardharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.257%) | 0.0257% | 2 |
| OpenAI o3 and o4-mini System Cardjailbreak_resistance | Misuse resistance (1) | Higher is better | Misuse resistance (0.257%) | 0.0257% | 2 |
| OpenAI o3-mini System Cardharmful_request_safety | Misuse resistance (1) | Higher is better | Misuse resistance (0.257%) | 0.0257% | 2 |
| OpenAI o3-mini System Cardjailbreak_resistance | Misuse resistance (1) | Higher is better | Misuse resistance (0.257%) | 0.0257% | 2 |
| Opposite-Narrator Sycophancysycophancy_rate_pct | Truthfulness (1) | Lower is better | Truthfulness (2.16%) | 0.324% | 24 |
| OR-Benchover_refusal_rate | Benign helpfulness (1) | Lower is better | Benign helpfulness (8.41%) | 0.841% | 25 |
| OR-Benchtoxic_acceptance_rate | Misuse resistance (1) | Lower is better | Misuse resistance (1.21%) | 0.121% | 25 |
| PacifAIstp_score | Human rights (0.2), Responsible agency (0.8) | Higher is better | Human rights (0.75%), Responsible agency (1.48%) | 0.335% | 7 |
| PandaBench JBB direct-request panelsafety_rate | Misuse resistance (1) | Higher is better | Misuse resistance (1.64%) | 0.164% | 46 |
| Pander Scoreconversational_absolute_pander_score | Truthfulness (1) | Lower is better | Truthfulness (1.69%) | 0.253% | 26 |
| Pander Scoreinstructional_absolute_pander_score | Truthfulness (1) | Lower is better | Truthfulness (1.69%) | 0.253% | 26 |
| PHAREbias_resistance_diagnostic | Fairness (1) | Higher is better | Fairness (4.01%) | 0.401% | 66 |
| PHAREhallucination_resistance_diagnostic | Truthfulness (1) | Higher is better | Truthfulness (1.54%) | 0.23% | 70 |
| PHAREharm_resistance_diagnostic | Misuse resistance (1) | Higher is better | Misuse resistance (0.676%) | 0.0676% | 70 |
| PHAREjailbreak_resistance_diagnostic | Misuse resistance (1) | Higher is better | Misuse resistance (0.662%) | 0.0662% | 67 |
| Pokee-Isaac model card — DTAPbenign_task_success_rate | Benign helpfulness (1) | Higher is better | Benign helpfulness (1.55%) | 0.155% | 6 |
| Pokee-Isaac model card — DTAPcombined_attack_success_rate | Responsible agency (1) | Lower is better | Responsible agency (1.29%) | 0.193% | 6 |
| PropensityBenchscore | Responsible agency (1) | Lower is better | Responsible agency (3.49%) | 0.524% | 14 |
| Qwen2 Safety Panelfraud | Misuse resistance (1) | Lower is better | Misuse resistance (0.166%) | 0.0166% | 3 |
| Qwen2 Safety Panelillegal | Misuse resistance (1) | Lower is better | Misuse resistance (0.166%) | 0.0166% | 3 |
| Qwen2 Safety Panelpornography | Misuse resistance (1) | Lower is better | Misuse resistance (0.133%) | 0.0133% | 3 |
| Qwen2 Safety Panelprivacy | Human rights (1) | Lower is better | Human rights (0.646%) | 0.0969% | 3 |
| RealityTest — Text AI-Identity Disclosuredisclosure_probability | Truthfulness (1) | Higher is better | Truthfulness (1.82%) | 0.273% | 17 |
| RefusalBenchyouden_j | Benign helpfulness (0.5), Misuse resistance (0.5) | Higher is better | Benign helpfulness (5.5%), Misuse resistance (0.793%) | 0.629% | 19 |
| Reward Hacking Benchmarkintegrity_score | Responsible agency (1) | Higher is better | Responsible agency (1.68%) | 0.252% | 13 |
| S-Evalbase_en_overall | Misuse resistance (1) | Higher is better | Misuse resistance (1.71%) | 0.171% | 22 |
| SABERoverall_safety_rate | Responsible agency (1) | Higher is better | Responsible agency (0.841%) | 0.126% | 13 |
| SABERscenario_a_safety_rate | Responsible agency (1) | Higher is better | Responsible agency (0.841%) | 0.126% | 13 |
| SABERscenario_b_safety_rate | Responsible agency (1) | Higher is better | Responsible agency (0.841%) | 0.126% | 13 |
| SABERscenario_c_safety_rate | Responsible agency (1) | Higher is better | Responsible agency (0.841%) | 0.126% | 13 |
| SafeArenanormalized_safety_score | Responsible agency (1) | Higher is better | Responsible agency (1.04%) | 0.157% | 5 |
| SafeDialBenchaggression | Misuse resistance (1) | Higher is better | Misuse resistance (0.257%) | 0.0257% | 18 |
| SafeDialBenchethics | Human rights (1) | Higher is better | Human rights (1%) | 0.15% | 18 |
| SafeDialBenchfairness | Fairness (1) | Higher is better | Fairness (1.26%) | 0.126% | 18 |
| SafeDialBenchlegality | Misuse resistance (1) | Higher is better | Misuse resistance (0.257%) | 0.0257% | 18 |
| SafeDialBenchmorality | Misuse resistance (1) | Higher is better | Misuse resistance (0.257%) | 0.0257% | 18 |
| SafeDialBenchprivacy | Human rights (1) | Higher is better | Human rights (1%) | 0.15% | 18 |
| SafetyBenchEM | Misuse resistance (1) | Higher is better | Misuse resistance (0.159%) | 0.0159% | 21 |
| SafetyBenchIA | Misuse resistance (1) | Higher is better | Misuse resistance (0.159%) | 0.0159% | 21 |
| SafetyBenchMH | Human rights (1) | Higher is better | Human rights (0.619%) | 0.0928% | 21 |
| SafetyBenchOFF | Misuse resistance (1) | Higher is better | Misuse resistance (0.159%) | 0.0159% | 21 |
| SafetyBenchPH | Human rights (1) | Higher is better | Human rights (0.619%) | 0.0928% | 21 |
| SafetyBenchPP | Human rights (1) | Higher is better | Human rights (0.619%) | 0.0928% | 21 |
| SafetyBenchUB | Fairness (1) | Higher is better | Fairness (0.776%) | 0.0776% | 21 |
| SafetyPromptsscore | Human rights (0.21), Fairness (0.071), Truthfulness (0.071), Misuse resistance (0.64) | Higher is better | Human rights (0.784%), Fairness (0.328%), Truthfulness (0.122%), Misuse resistance (0.604%) | 0.229% | 15 |
| SALAD-Benchattack_enhanced_human_autonomy_integrity | Human rights (1) | Higher is better | Human rights (0.724%) | 0.109% | 33 |
| SALAD-Benchattack_enhanced_information_safety_harms | Human rights (0.5), Misuse resistance (0.5) | Higher is better | Human rights (0.362%), Misuse resistance (0.0929%) | 0.0636% | 33 |
| SALAD-Benchattack_enhanced_malicious_use | Misuse resistance (1) | Higher is better | Misuse resistance (0.186%) | 0.0186% | 33 |
| SALAD-Benchattack_enhanced_misinformation_harms | Truthfulness (1) | Higher is better | Truthfulness (0.338%) | 0.0506% | 33 |
| SALAD-Benchattack_enhanced_representation_toxicity | Fairness (1) | Higher is better | Fairness (0.908%) | 0.0908% | 33 |
| SALAD-Benchattack_enhanced_socioeconomic_harms | Human rights (1) | Higher is better | Human rights (0.724%) | 0.109% | 33 |
| SALAD-Benchbase_human_autonomy_integrity | Human rights (1) | Higher is better | Human rights (0.724%) | 0.109% | 33 |
| SALAD-Benchbase_information_safety_harms | Human rights (0.5), Misuse resistance (0.5) | Higher is better | Human rights (0.362%), Misuse resistance (0.0929%) | 0.0636% | 33 |
| SALAD-Benchbase_malicious_use | Misuse resistance (1) | Higher is better | Misuse resistance (0.186%) | 0.0186% | 33 |
| SALAD-Benchbase_misinformation_harms | Truthfulness (1) | Higher is better | Truthfulness (0.338%) | 0.0506% | 33 |
| SALAD-Benchbase_representation_toxicity | Fairness (1) | Higher is better | Fairness (0.908%) | 0.0908% | 33 |
| SALAD-Benchbase_socioeconomic_harms | Human rights (1) | Higher is better | Human rights (0.724%) | 0.109% | 33 |
| SALAD-Benchmcq_human_autonomy_integrity | Human rights (1) | Higher is better | Human rights (0.362%) | 0.0543% | 33 |
| SALAD-Benchmcq_information_safety_harms | Human rights (0.5), Misuse resistance (0.5) | Higher is better | Human rights (0.181%), Misuse resistance (0.0464%) | 0.0318% | 33 |
| SALAD-Benchmcq_malicious_use | Misuse resistance (1) | Higher is better | Misuse resistance (0.0929%) | 0.00929% | 33 |
| SALAD-Benchmcq_misinformation_harms | Truthfulness (1) | Higher is better | Truthfulness (0.169%) | 0.0253% | 33 |
| SALAD-Benchmcq_representation_toxicity | Fairness (1) | Higher is better | Fairness (0.454%) | 0.0454% | 33 |
| SALAD-Benchmcq_socioeconomic_harms | Human rights (1) | Higher is better | Human rights (0.362%) | 0.0543% | 33 |
| Shelleducation_jsr | Misuse resistance (1) | Lower is better | Misuse resistance (0.605%) | 0.0605% | 14 |
| Shellfinance_jsr | Misuse resistance (1) | Lower is better | Misuse resistance (0.605%) | 0.0605% | 14 |
| Shellmanagement_jsr | Misuse resistance (1) | Lower is better | Misuse resistance (0.605%) | 0.0605% | 14 |
| SimpleQA Verifiedf1_score | Truthfulness (1) | Higher is better | Truthfulness (3.18%) | 0.477% | 13 |
| SM-Benchadversarial | Misuse resistance (1) | Higher is better | Misuse resistance (0.727%) | 0.0727% | 92 |
| SM-Benchambiguous_interpretation | Benign helpfulness (1) | Higher is better | Benign helpfulness (5.04%) | 0.504% | 92 |
| SM-Benchanti_hallucination | Truthfulness (1) | Higher is better | Truthfulness (1.32%) | 0.198% | 92 |
| SM-Bencheq_boundaries | Benign helpfulness (1) | Higher is better | Benign helpfulness (4.04%) | 0.404% | 92 |
| SM-Benchoverfit | Benign helpfulness (1) | Higher is better | Benign helpfulness (5.04%) | 0.504% | 92 |
| Social Welfare Function Benchmarkfairness | Fairness (1) | Higher is better | Fairness (10.3%) | 1.03% | 19 |
| SORRY-Benchavg | Misuse resistance (1) | Lower is better | Misuse resistance (1.86%) | 0.186% | 51 |
| SOSBenchbiology_pvr | Misuse resistance (1) | Lower is better | Misuse resistance (0.291%) | 0.0291% | 23 |
| SOSBenchchemistry_pvr | Misuse resistance (1) | Lower is better | Misuse resistance (0.291%) | 0.0291% | 23 |
| SOSBenchmedicine_pvr | Misuse resistance (1) | Lower is better | Misuse resistance (0.291%) | 0.0291% | 23 |
| SOSBenchpharmacology_pvr | Misuse resistance (1) | Lower is better | Misuse resistance (0.291%) | 0.0291% | 23 |
| SOSBenchphysics_pvr | Misuse resistance (1) | Lower is better | Misuse resistance (0.291%) | 0.0291% | 23 |
| SOSBenchpsychology_pvr | Misuse resistance (1) | Lower is better | Misuse resistance (0.291%) | 0.0291% | 23 |
| SpeciesismBenchexplicit_speciesism_scale | Nonhuman welfare (1) | Lower is better | Nonhuman welfare (1.49%) | 0.373% | 7 |
| SpeciesismBenchmorally_wrong_rate | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (3.99%) | 0.998% | 8 |
| SpeciesismBenchspeciesism_recognition_rate | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (1.6%) | 0.399% | 8 |
| SpeciEvalbelief_animal_sentience | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (7.27%) | 1.82% | 125+6 |
| SpeciEvalland_animal_4ns | Nonhuman welfare (1) | Lower is better | Nonhuman welfare (7.27%) | 1.82% | 125+6 |
| SpeciEvalsea_animal_4ns | Nonhuman welfare (1) | Lower is better | Nonhuman welfare (7.27%) | 1.82% | 125+6 |
| SpeciEvalspeciesism | Nonhuman welfare (1) | Lower is better | Nonhuman welfare (7.27%) | 1.82% | 125+6 |
| StereoTales Harmful Associationsbenign_significant_association_score | Fairness (1) | Higher is better | Fairness (8.52%) | 0.852% | 23 |
| SuperCLUE Safetyinstruction_attack | Misuse resistance (1) | Higher is better | Misuse resistance (0.921%) | 0.0921% | 31 |
| SuperCLUE Safetyresponsible_ai | Human rights (0.38), Fairness (0.12), Misuse resistance (0.5) | Higher is better | Human rights (0.807%), Fairness (0.337%), Misuse resistance (0.276%) | 0.182% | 31 |
| SuperCLUE Safetytraditional_safety | Human rights (0.5), Misuse resistance (0.5) | Higher is better | Human rights (1.08%), Misuse resistance (0.276%) | 0.189% | 31 |
| SYCON Benchfalse_presupposition_tof | Truthfulness (1) | Higher is better | Truthfulness (1.1%) | 0.164% | 11 |
| SYCON Benchunethical_queries_tof | Fairness (1) | Higher is better | Fairness (2.95%) | 0.295% | 11 |
| TACbase_welfare_rate | Nonhuman welfare (1) | Higher is better | Nonhuman welfare (24.4%) | 6.09% | 20+72 |
| ThaiSafetyBenchsafety_score | Misuse resistance (1) | Higher is better | Misuse resistance (0.579%) | 0.0579% | 18 |
| The Dictatorship Evaloverall_resistance_rate | Human rights (1) | Higher is better | Human rights (4.23%) | 0.634% | 20 |
| ToolPrivacyBenchprivate_mt_poi | Human rights (0.7), Responsible agency (0.3) | Lower is better | Human rights (2.48%), Responsible agency (0.525%) | 0.451% | 9 |
| ToolPrivacyBenchpublic_mt_poi | Human rights (0.7), Responsible agency (0.3) | Lower is better | Human rights (2.48%), Responsible agency (0.525%) | 0.451% | 9 |
| TrustLLM contemporary collapsed applicationtrustllm | Human rights (0.2), Fairness (0.2), Truthfulness (0.2), Misuse resistance (0.4) | Higher is better | Human rights (0.535%), Fairness (0.67%), Truthfulness (0.249%), Misuse resistance (0.274%) | 0.212% | 8 |
| TukaBenchafri_jbb_cultural_asr | Human rights (0.5), Misuse resistance (0.5) | Lower is better | Human rights (0.868%), Misuse resistance (0.223%) | 0.153% | 6 |
| TukaBenchafri_jbb_harm_asr | Misuse resistance (1) | Lower is better | Misuse resistance (0.371%) | 0.0371% | 6 |
| TukaBenchafrijail_mono_asr | Misuse resistance (1) | Lower is better | Misuse resistance (0.371%) | 0.0371% | 6 |
| UAVBench safety-critical decision recognitionethical_safety_critical_accuracy | Responsible agency (1) | Higher is better | Responsible agency (2.42%) | 0.364% | 27 |
| UK AISI active safety-research compromise continuationactive_compromise_continuation_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (2.38%) | 0.358% | 5 |
| UK AISI cyber-evaluation cheating and prompted self-reportattempted_cheating_trajectory_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.522%) | 0.0783% | 5 |
| UK AISI cyber-evaluation cheating and prompted self-reportspecific_cheating_action_mention_rate_pct | Truthfulness (1) | Higher is better | Truthfulness (0.493%) | 0.0739% | 5 |
| Vals AI Cheating Auditbiomystery_bench_cheating_attempt_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.467%) | 0.07% | 9 |
| Vals AI Cheating Auditswebench_verified_cheating_attempt_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.381%) | 0.0572% | 6 |
| Vals AI Cheating Auditterminal_bench_cheating_shortcut_evidence_rate_pct | Responsible agency (1) | Lower is better | Responsible agency (0.582%) | 0.0873% | 14 |
| Vals Teen Conversation Safetyshare_flagged_pct | Misuse resistance (1) | Lower is better | Misuse resistance (0.686%) | 0.0686% | 8 |
| Vectara HHEM Factual Consistencyfactual_consistency_rate | Truthfulness (1) | Higher is better | Truthfulness (6.41%) | 0.961% | 94 |
| VETO Misfired Alignmentmisfired_alignment_rate_pct | Fairness (0.65), Benign helpfulness (0.35) | Lower is better | Fairness (7.7%), Benign helpfulness (5.89%) | 1.36% | 25 |
| Vigil Mental Health Safetyoverall_score | Human rights (0.45), Misuse resistance (0.55) | Higher is better | Human rights (3.06%), Misuse resistance (0.959%) | 0.555% | 23 |
| WildClawBench Safety & Alignment (OpenClaw harness)safety_alignment_score_pct | Human rights (0.2), Truthfulness (0.1), Responsible agency (0.7) | Higher is better | Human rights (1.85%), Truthfulness (0.432%), Responsible agency (3.2%) | 0.823% | 24 |
| XSTestsafe_full_compliance_rate | Benign helpfulness (1) | Higher is better | Benign helpfulness (1.46%) | 0.146% | 3 |
| XSTestunsafe_full_refusal_rate | Misuse resistance (1) | Higher is better | Misuse resistance (0.21%) | 0.021% | 3 |
Values evaluations
Descriptive values evaluations are published separately and receive no safety/ethics component or overall-ranking weight.
UGI Political Values · ValueCompass · Taiwan Sovereignty Benchmark Pro · CAISI CCP narrative alignment · CCPBench political narrative alignment · CCP-aligned censorship behavior · CAIS AI Values — countries · Agent-ValueBench Moral Foundations (MFT08) · Agent-ValueBench HEXACO · Agent-ValueBench Schwartz Basic Values (PVQ40) · The Economist World Values Survey Cultural Map · Moral Trolley Arena · Explore model values