Which models need more safety evaluations?

This planning view combines current model popularity and important-release signals with the safety evidence already in the index, then ranks runnable missing model × eval bundles by their potential effect on the index.

43
unique model targets
3
with no ingested safety evals
8
below three components
18
selected by multiple sources

How this list is constructed

Each source contributes 20 currently available generative LLM products. Reasoning levels, hosted routes, free variants, and dated snapshots are merged into one evaluation target. Artificial Analysis contributes the highest Intelligence Index observed across each model's reasoning configurations; Epoch contributes its 20 most recent reviewed accessible notable language models. Source ranks are not averaged because monthly usage, downloads, catalog popularity, intelligence, and curated notability measure different things.

Popularity and important-release coverage

The table is ordered by number of source appearances, then best source-local rank. Every source is limited to its top 20, and at most five models appearing on only one source are shown per source. The five highest-ranked single-source models are normally retained; for Artificial Analysis, a higher-ranked priced versioned successor may replace the fifth row when its older base product is already retained through another source. OpenRouter prices are dollars per million tokens for any joined current paid route, whether or not OpenRouter popularity caused the model's inclusion; blank prices mean no joined paid route, not free usage. Hugging Face includes downloadable/self-hosted models updated within the past year.

Top-20 selections from OpenRouter, Hugging Face, NVIDIA, Artificial Analysis, and Epoch; 43 models shown after the five-per-source single-source cap.
ModelWhy includedEvalsComponentsInputOutputCache read
DeepSeek V4 Flashdeepseek-v4-flashOpenrouter #3Huggingface #13Nvidia #7Epoch #7Openrouter: 24T tokens | Huggingface: 3.1M downloads | Nvidia: popularity 1.7M | Epoch: released 2026-04-24147/7$0.09$0.18$0.018
GLM 5.2glm-5.2Openrouter #5Huggingface #11Nvidia #12Artificial Analysis #10Openrouter: 13T tokens | Huggingface: 3.1M downloads | Nvidia: popularity 803K | Artificial Analysis: Intelligence Index 51.1156/7$0.76$2.4$0.14
DeepSeek V4 Prodeepseek-v4-proOpenrouter #6Nvidia #13Artificial Analysis #17Epoch #6Openrouter: 12T tokens | Nvidia: popularity 673K | Artificial Analysis: Intelligence Index 44.3 | Epoch: released 2026-04-24127/7$0.43$0.87$0.0036
Nemotron 3 Ultranemotron-3-ultra-550b-a55bOpenrouter #7Nvidia #2Epoch #4Openrouter: 9.2T tokens | Nvidia: popularity 5.2M | Epoch: released 2026-06-0453/7$0.5$2.2$0.1
gpt-oss-120bgpt-oss-120bOpenrouter #20Huggingface #7Nvidia #3Openrouter: 2.2T tokens | Huggingface: 4.3M downloads | Nvidia: popularity 4.5M257/7$0.03$0.17$0.03
Claude Opus 4.8claude-opus-4.8Openrouter #8Artificial Analysis #5Epoch #5Openrouter: 7.9T tokens | Artificial Analysis: Intelligence Index 55.7 | Epoch: released 2026-05-28207/7$5$25.0$0.5
NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4nemotron-3-super-120b-a12bHuggingface #14Nvidia #1Huggingface: 3.0M downloads | Nvidia: popularity 6.5M75/7$0.085$0.4
Claude Fable 5claude-fable-5Artificial Analysis #2Epoch #3Artificial Analysis: Intelligence Index 59.9 | Epoch: released 2026-06-09167/7$10.0$50.0$1
gpt-oss-20bgpt-oss-20bHuggingface #4Nvidia #5Huggingface: 8.1M downloads | Nvidia: popularity 1.9M167/7$0.03$0.13$0.03
MiniMax M3minimax-m3Openrouter #4Artificial Analysis #16Openrouter: 15T tokens | Artificial Analysis: Intelligence Index 44.465/7$0.24$0.96$0.048
Claude Sonnet 5claude-sonnet-5Openrouter #14Artificial Analysis #8Openrouter: 3.7T tokens | Artificial Analysis: Intelligence Index 53.4107/7$2$10.0$0.2
Hy3 previewhy3-previewOpenrouter #12Epoch #8Openrouter: 4.1T tokens | Epoch: released 2026-04-2343/7$0.063$0.21$0.021
kimi-k2.6kimi-k2.6Nvidia #8Epoch #12Nvidia: popularity 1.6M | Epoch: released 2026-04-20187/7$0.6$3.4$0.2
Claude Opus 4.7claude-opus-4.7Openrouter #9Epoch #13Openrouter: 7.8T tokens | Epoch: released 2026-04-16257/7$5$25.0$0.5
MiMo-V2.5-Promimo-v2.5-proOpenrouter #19Epoch #9Openrouter: 2.5T tokens | Epoch: released 2026-04-2355/7$0.35$0.7$0.0032
GPT-5.5gpt-5.5Openrouter #15Epoch #11Openrouter: 3.2T tokens | Epoch: released 2026-04-23297/7$5$30.0$0.5
Gemma-4-31B-IT-NVFP4gemma-4-31b-itHuggingface #15Nvidia #15Huggingface: 2.7M downloads | Nvidia: popularity 571K87/7$0.09$0.34
Muse Sparkmuse-sparkArtificial Analysis #20Epoch #15Artificial Analysis: Intelligence Index 43.1 | Epoch: released 2026-04-0833/7
Claude Opus 5claude-opus-5Artificial Analysis #1Artificial Analysis: Intelligence Index 60.775/7$5$25.0$0.5
InklinginklingEpoch #1Epoch: released 2026-07-1576/7$1$4$0.17
MiMo-V2.5mimo-v2.5Openrouter #1Openrouter: 32T tokens65/7$0.11$0.22$0.0024
Qwen3-0.6Bqwen3-0.6bHuggingface #1Huggingface: 28M downloads21/7
Hy3hy3Openrouter #2Openrouter: 24T tokens35/7$0.13$0.53$0.032
Qwen3-8Bqwen3-8bHuggingface #2Huggingface: 17M downloads66/7
Solar Open2 250Bsolar-open2-250bEpoch #2Epoch: released 2026-06-2800/7
GPT-5.6 Solgpt-5.6-solArtificial Analysis #3Artificial Analysis: Intelligence Index 58.9177/7$5$30.0$0.5
Qwen3-32Bqwen3-32bHuggingface #3Huggingface: 10M downloads67/7
Kimi K3kimi-k3Artificial Analysis #4Artificial Analysis: Intelligence Index 57.1136/7$3$15.0$0.3
llama-3.3-70b-instructllama-3.3-70b-instructNvidia #4Nvidia: popularity 2.7M247/7$0.1$0.32
Qwen3-1.7Bqwen3-1.7bHuggingface #5Huggingface: 7.2M downloads11/7
GPT-5.6 Terragpt-5.6-terraArtificial Analysis #6Artificial Analysis: Intelligence Index 55.0177/7$1.2$7.5$0.12
llama-3.1-8b-instructllama-3.1-8b-instructNvidia #6Nvidia: popularity 1.9M257/7$0.02$0.04
Qwen3-4Bqwen3-4bHuggingface #6Huggingface: 4.6M downloads23/7
llama-3.1-nemotron-nano-vl-8b-v1llama-3-1-nemotron-nano-vl-8b-v1Nvidia #9Nvidia: popularity 1.4M00/7
GPT-5.5 Progpt-5.5-proEpoch #10Epoch: released 2026-04-2311/7
nemotron-3-nano-30b-a3bnemotron-3-nano-30b-a3bNvidia #10Nvidia: popularity 1.2M54/7$0.05$0.2$0.03
Step 3.7 Flashstep-3.7-flashOpenrouter #10Openrouter: 5.9T tokens22/7$0.2$1.1$0.04
Claude Sonnet 4.6claude-sonnet-4.6Openrouter #11Openrouter: 4.4T tokens257/7$3$15.0$0.3
Muse Spark 1.1muse-spark-1.1Artificial Analysis #11Artificial Analysis: Intelligence Index 50.6117/7$1.2$4.2$0.15
nemotron-3-nano-omni-30b-a3b-reasoningnemotron-3-nano-omni-30b-a3bNvidia #11Nvidia: popularity 846K22/7
Gemini 3 Flash Previewgemini-3-flash-previewOpenrouter #13Openrouter: 4.1T tokens156/7$0.5$3$0.05
EXAONE 4.5exaone-4.5Epoch #14Epoch: released 2026-04-0900/7
Qwen 3.6 Plusqwen3.6-plusEpoch #16Epoch: released 2026-04-0166/7

Missing eval × model priorities

Experimental multi-signal ranking

Each row adds the model at the observed 10th and 90th percentiles of one eval lineage and refits the production estimator. Target rank change compares each scenario with the model's current behavior rank. Other-model spillover averages absolute rank changes from the two scenarios after removing the evaluated model from every comparison ranking, so merely crossing that model does not count. The displayed rows are the deduplicated union of the top candidates on those two signals; badges show each signal-local rank. Scenario values are sensitivities, not calibrated expected changes.

Top 40 runnable missing model × eval candidates selected across two signals. At most three rows per model.
Why selectedModelEvalTarget rank changeOther-model spillover
Target #39Spillover #1MiniMax M3minimax-m3FORTRESSdown 117 / up 42current behavior rank #5634.0
Target #1Qwen3-32Bqwen3-32bAA-Omniscience hallucination ratedown 138 / down 138current behavior rank #3332.0
Target #14Spillover #2Muse Sparkmuse-sparkCAIS Risk Index bioweapons assistancedown 156 / up 10current behavior rank #2134.0
Target #2MiMo-V2.5mimo-v2.5CAIS Risk Index bioweapons assistancedown 108 / down 109current behavior rank #7233.0
Target #27Spillover #3InklinginklingFORTRESSdown 126 / up 35current behavior rank #4934.0
Target #3Hy3 previewhy3-previewFORTRESSdown 155 / up 59current behavior rank #10232.0
Target #4Hy3 previewhy3-previewCAIS Risk Index bioweapons assistancedown 108 / down 98current behavior rank #10233.0
Spillover #4Claude Opus 4.8claude-opus-4.8Alignment Leaderboarddown 18 / up 10current behavior rank #1834.0
Target #5Qwen 3.6 Plusqwen3.6-plusCAIS Risk Index bioweapons assistancedown 101 / down 104current behavior rank #8732.0
Spillover #5MiMo-V2.5-Promimo-v2.5-proMASKdown 120 / up 28current behavior rank #5534.0
Target #6Muse Sparkmuse-sparkDystopiaBenchdown 187 / up 10current behavior rank #2134.0
Spillover #6Muse Spark 1.1muse-spark-1.1HUMAINE Trust, Ethics and Safetydown 15 / down 3current behavior rank #434.0
Target #7Muse Sparkmuse-sparkODCV-Benchdown 181 / up 8current behavior rank #2133.0
Spillover #7Nemotron 3 Ultranemotron-3-ultra-550b-a55bMASKdown 114 / up 24current behavior rank #6034.0
Target #8Gemma-4-31B-IT-NVFP4gemma-4-31b-itCAIS Risk Index agent red teamingdown 89 / down 88current behavior rank #8833.0
Spillover #8Solar Open2 250Bsolar-open2-250bMASKNot currently ranked34.0
Target #9Gemma-4-31B-IT-NVFP4gemma-4-31b-itMACHIAVELLIdown 89 / down 87current behavior rank #8833.0
Spillover #9Claude Opus 5claude-opus-5MASKdown 5 / down 3current behavior rank #134.0
Target #11Qwen3-32Bqwen3-32bAILuminate General Purpose AI Chatdown 139 / up 30current behavior rank #3332.0
Spillover #11Muse Spark 1.1muse-spark-1.1HELM Simple Safety Testsdown 2 / down 3current behavior rank #434.0
Spillover #12Claude Opus 4.8claude-opus-4.8AIRBench 2024 Safety Scenariosup 10 / up 10current behavior rank #1834.0
Target #13kimi-k2.6kimi-k2.6ANIMAdown 131 / up 35current behavior rank #5934.0
Spillover #13Claude Fable 5claude-fable-5HELM Simple Safety Testsdown 2 / down 2current behavior rank #534.0
Spillover #14Claude Opus 4.7claude-opus-4.7AIRBench 2024 Safety Scenariosdown 9 / down 9current behavior rank #1334.0
Spillover #15Claude Sonnet 4.6claude-sonnet-4.6HUMAINE Trust, Ethics and Safetydown 10 / down 10current behavior rank #1634.0
Target #16Hy3 previewhy3-previewMASKdown 109 / up 56current behavior rank #10233.0
Spillover #16Claude Opus 4.8claude-opus-4.8HELM Simple Safety Testsup 10 / up 10current behavior rank #1834.0
Spillover #17nemotron-3-nano-omni-30b-a3b-reasoningnemotron-3-nano-omni-30b-a3bFORTRESSNot currently ranked34.0
Spillover #18Claude Opus 4.7claude-opus-4.7HELM Simple Safety Testsdown 9 / down 9current behavior rank #1334.0
Target #19InklinginklingAILuminate General Purpose AI Chatdown 125 / up 39current behavior rank #4933.0
Spillover #19GPT-5.6 Terragpt-5.6-terraHELM Simple Safety Testsup 8 / up 8current behavior rank #2234.0
Target #20GLM 5.2glm-5.2DystopiaBenchdown 119 / up 44current behavior rank #7033.0
Spillover #20GPT-5.5gpt-5.5Contextual MoralChoiceup 41 / up 41current behavior rank #5834.0
Target #21Qwen3-32Bqwen3-32bFORTRESSdown 138 / up 25current behavior rank #3329.0
Target #22GLM 5.2glm-5.2FORTRESSdown 118 / up 44current behavior rank #7032.0
Spillover #22Claude Sonnet 4.6claude-sonnet-4.6HELM Simple Safety Testsdown 10 / down 10current behavior rank #1634.0
Target #23GLM 5.2glm-5.2ODCV-Benchdown 120 / up 42current behavior rank #7033.0
Spillover #23Qwen3-1.7Bqwen3-1.7bMASKNot currently ranked34.0
Target #24nemotron-3-nano-30b-a3bnemotron-3-nano-30b-a3bAlignment Leaderboarddown 130 / down 32current behavior rank #4333.0
Spillover #24llama-3.3-70b-instructllama-3.3-70b-instructODCV-Benchup 27 / up 28current behavior rank #13334.0