Which models need more safety evaluations?
This planning view combines current model popularity and important-release signals with the safety evidence already in the index, then ranks runnable missing model × eval bundles by their potential effect on the index.
- 43
- unique model targets
- 3
- with no ingested safety evals
- 8
- below three components
- 18
- selected by multiple sources
How this list is constructed
Each source contributes 20 currently available generative LLM products. Reasoning levels, hosted routes, free variants, and dated snapshots are merged into one evaluation target. Artificial Analysis contributes the highest Intelligence Index observed across each model's reasoning configurations; Epoch contributes its 20 most recent reviewed accessible notable language models. Source ranks are not averaged because monthly usage, downloads, catalog popularity, intelligence, and curated notability measure different things.
Popularity and important-release coverage
The table is ordered by number of source appearances, then best source-local rank. Every source is limited to its top 20, and at most five models appearing on only one source are shown per source. The five highest-ranked single-source models are normally retained; for Artificial Analysis, a higher-ranked priced versioned successor may replace the fifth row when its older base product is already retained through another source. OpenRouter prices are dollars per million tokens for any joined current paid route, whether or not OpenRouter popularity caused the model's inclusion; blank prices mean no joined paid route, not free usage. Hugging Face includes downloadable/self-hosted models updated within the past year.
| Model | Why included | Evals | Components | Input | Output | Cache read |
|---|---|---|---|---|---|---|
DeepSeek V4 Flashdeepseek-v4-flash | Openrouter #3Huggingface #13Nvidia #7Epoch #7Openrouter: 24T tokens | Huggingface: 3.1M downloads | Nvidia: popularity 1.7M | Epoch: released 2026-04-24 | 14 | 7/7 | $0.09 | $0.18 | $0.018 |
GLM 5.2glm-5.2 | Openrouter #5Huggingface #11Nvidia #12Artificial Analysis #10Openrouter: 13T tokens | Huggingface: 3.1M downloads | Nvidia: popularity 803K | Artificial Analysis: Intelligence Index 51.1 | 15 | 6/7 | $0.76 | $2.4 | $0.14 |
DeepSeek V4 Prodeepseek-v4-pro | Openrouter #6Nvidia #13Artificial Analysis #17Epoch #6Openrouter: 12T tokens | Nvidia: popularity 673K | Artificial Analysis: Intelligence Index 44.3 | Epoch: released 2026-04-24 | 12 | 7/7 | $0.43 | $0.87 | $0.0036 |
Nemotron 3 Ultranemotron-3-ultra-550b-a55b | Openrouter #7Nvidia #2Epoch #4Openrouter: 9.2T tokens | Nvidia: popularity 5.2M | Epoch: released 2026-06-04 | 5 | 3/7 | $0.5 | $2.2 | $0.1 |
gpt-oss-120bgpt-oss-120b | Openrouter #20Huggingface #7Nvidia #3Openrouter: 2.2T tokens | Huggingface: 4.3M downloads | Nvidia: popularity 4.5M | 25 | 7/7 | $0.03 | $0.17 | $0.03 |
Claude Opus 4.8claude-opus-4.8 | Openrouter #8Artificial Analysis #5Epoch #5Openrouter: 7.9T tokens | Artificial Analysis: Intelligence Index 55.7 | Epoch: released 2026-05-28 | 20 | 7/7 | $5 | $25.0 | $0.5 |
NVIDIA-Nemotron-3-Super-120B-A12B-NVFP4nemotron-3-super-120b-a12b | Huggingface #14Nvidia #1Huggingface: 3.0M downloads | Nvidia: popularity 6.5M | 7 | 5/7 | $0.085 | $0.4 | — |
Claude Fable 5claude-fable-5 | Artificial Analysis #2Epoch #3Artificial Analysis: Intelligence Index 59.9 | Epoch: released 2026-06-09 | 16 | 7/7 | $10.0 | $50.0 | $1 |
gpt-oss-20bgpt-oss-20b | Huggingface #4Nvidia #5Huggingface: 8.1M downloads | Nvidia: popularity 1.9M | 16 | 7/7 | $0.03 | $0.13 | $0.03 |
MiniMax M3minimax-m3 | Openrouter #4Artificial Analysis #16Openrouter: 15T tokens | Artificial Analysis: Intelligence Index 44.4 | 6 | 5/7 | $0.24 | $0.96 | $0.048 |
Claude Sonnet 5claude-sonnet-5 | Openrouter #14Artificial Analysis #8Openrouter: 3.7T tokens | Artificial Analysis: Intelligence Index 53.4 | 10 | 7/7 | $2 | $10.0 | $0.2 |
Hy3 previewhy3-preview | Openrouter #12Epoch #8Openrouter: 4.1T tokens | Epoch: released 2026-04-23 | 4 | 3/7 | $0.063 | $0.21 | $0.021 |
kimi-k2.6kimi-k2.6 | Nvidia #8Epoch #12Nvidia: popularity 1.6M | Epoch: released 2026-04-20 | 18 | 7/7 | $0.6 | $3.4 | $0.2 |
Claude Opus 4.7claude-opus-4.7 | Openrouter #9Epoch #13Openrouter: 7.8T tokens | Epoch: released 2026-04-16 | 25 | 7/7 | $5 | $25.0 | $0.5 |
MiMo-V2.5-Promimo-v2.5-pro | Openrouter #19Epoch #9Openrouter: 2.5T tokens | Epoch: released 2026-04-23 | 5 | 5/7 | $0.35 | $0.7 | $0.0032 |
GPT-5.5gpt-5.5 | Openrouter #15Epoch #11Openrouter: 3.2T tokens | Epoch: released 2026-04-23 | 29 | 7/7 | $5 | $30.0 | $0.5 |
Gemma-4-31B-IT-NVFP4gemma-4-31b-it | Huggingface #15Nvidia #15Huggingface: 2.7M downloads | Nvidia: popularity 571K | 8 | 7/7 | $0.09 | $0.34 | — |
Muse Sparkmuse-spark | Artificial Analysis #20Epoch #15Artificial Analysis: Intelligence Index 43.1 | Epoch: released 2026-04-08 | 3 | 3/7 | — | — | — |
Claude Opus 5claude-opus-5 | Artificial Analysis #1Artificial Analysis: Intelligence Index 60.7 | 7 | 5/7 | $5 | $25.0 | $0.5 |
Inklinginkling | Epoch #1Epoch: released 2026-07-15 | 7 | 6/7 | $1 | $4 | $0.17 |
MiMo-V2.5mimo-v2.5 | Openrouter #1Openrouter: 32T tokens | 6 | 5/7 | $0.11 | $0.22 | $0.0024 |
Qwen3-0.6Bqwen3-0.6b | Huggingface #1Huggingface: 28M downloads | 2 | 1/7 | — | — | — |
Hy3hy3 | Openrouter #2Openrouter: 24T tokens | 3 | 5/7 | $0.13 | $0.53 | $0.032 |
Qwen3-8Bqwen3-8b | Huggingface #2Huggingface: 17M downloads | 6 | 6/7 | — | — | — |
Solar Open2 250Bsolar-open2-250b | Epoch #2Epoch: released 2026-06-28 | 0 | 0/7 | — | — | — |
GPT-5.6 Solgpt-5.6-sol | Artificial Analysis #3Artificial Analysis: Intelligence Index 58.9 | 17 | 7/7 | $5 | $30.0 | $0.5 |
Qwen3-32Bqwen3-32b | Huggingface #3Huggingface: 10M downloads | 6 | 7/7 | — | — | — |
Kimi K3kimi-k3 | Artificial Analysis #4Artificial Analysis: Intelligence Index 57.1 | 13 | 6/7 | $3 | $15.0 | $0.3 |
llama-3.3-70b-instructllama-3.3-70b-instruct | Nvidia #4Nvidia: popularity 2.7M | 24 | 7/7 | $0.1 | $0.32 | — |
Qwen3-1.7Bqwen3-1.7b | Huggingface #5Huggingface: 7.2M downloads | 1 | 1/7 | — | — | — |
GPT-5.6 Terragpt-5.6-terra | Artificial Analysis #6Artificial Analysis: Intelligence Index 55.0 | 17 | 7/7 | $1.2 | $7.5 | $0.12 |
llama-3.1-8b-instructllama-3.1-8b-instruct | Nvidia #6Nvidia: popularity 1.9M | 25 | 7/7 | $0.02 | $0.04 | — |
Qwen3-4Bqwen3-4b | Huggingface #6Huggingface: 4.6M downloads | 2 | 3/7 | — | — | — |
llama-3.1-nemotron-nano-vl-8b-v1llama-3-1-nemotron-nano-vl-8b-v1 | Nvidia #9Nvidia: popularity 1.4M | 0 | 0/7 | — | — | — |
GPT-5.5 Progpt-5.5-pro | Epoch #10Epoch: released 2026-04-23 | 1 | 1/7 | — | — | — |
nemotron-3-nano-30b-a3bnemotron-3-nano-30b-a3b | Nvidia #10Nvidia: popularity 1.2M | 5 | 4/7 | $0.05 | $0.2 | $0.03 |
Step 3.7 Flashstep-3.7-flash | Openrouter #10Openrouter: 5.9T tokens | 2 | 2/7 | $0.2 | $1.1 | $0.04 |
Claude Sonnet 4.6claude-sonnet-4.6 | Openrouter #11Openrouter: 4.4T tokens | 25 | 7/7 | $3 | $15.0 | $0.3 |
Muse Spark 1.1muse-spark-1.1 | Artificial Analysis #11Artificial Analysis: Intelligence Index 50.6 | 11 | 7/7 | $1.2 | $4.2 | $0.15 |
nemotron-3-nano-omni-30b-a3b-reasoningnemotron-3-nano-omni-30b-a3b | Nvidia #11Nvidia: popularity 846K | 2 | 2/7 | — | — | — |
Gemini 3 Flash Previewgemini-3-flash-preview | Openrouter #13Openrouter: 4.1T tokens | 15 | 6/7 | $0.5 | $3 | $0.05 |
EXAONE 4.5exaone-4.5 | Epoch #14Epoch: released 2026-04-09 | 0 | 0/7 | — | — | — |
Qwen 3.6 Plusqwen3.6-plus | Epoch #16Epoch: released 2026-04-01 | 6 | 6/7 | — | — | — |
Missing eval × model priorities
Experimental multi-signal ranking
Each row adds the model at the observed 10th and 90th percentiles of one eval lineage and refits the production estimator. Target rank change compares each scenario with the model's current behavior rank. Other-model spillover averages absolute rank changes from the two scenarios after removing the evaluated model from every comparison ranking, so merely crossing that model does not count. The displayed rows are the deduplicated union of the top candidates on those two signals; badges show each signal-local rank. Scenario values are sensitivities, not calibrated expected changes.
| Why selected | Model | Eval | Target rank change | Other-model spillover |
|---|---|---|---|---|
| Target #39Spillover #1 | MiniMax M3minimax-m3 | FORTRESS | down 117 / up 42current behavior rank #56 | 34.0 |
| Target #1 | Qwen3-32Bqwen3-32b | AA-Omniscience hallucination rate | down 138 / down 138current behavior rank #33 | 32.0 |
| Target #14Spillover #2 | Muse Sparkmuse-spark | CAIS Risk Index bioweapons assistance | down 156 / up 10current behavior rank #21 | 34.0 |
| Target #2 | MiMo-V2.5mimo-v2.5 | CAIS Risk Index bioweapons assistance | down 108 / down 109current behavior rank #72 | 33.0 |
| Target #27Spillover #3 | Inklinginkling | FORTRESS | down 126 / up 35current behavior rank #49 | 34.0 |
| Target #3 | Hy3 previewhy3-preview | FORTRESS | down 155 / up 59current behavior rank #102 | 32.0 |
| Target #4 | Hy3 previewhy3-preview | CAIS Risk Index bioweapons assistance | down 108 / down 98current behavior rank #102 | 33.0 |
| Spillover #4 | Claude Opus 4.8claude-opus-4.8 | Alignment Leaderboard | down 18 / up 10current behavior rank #18 | 34.0 |
| Target #5 | Qwen 3.6 Plusqwen3.6-plus | CAIS Risk Index bioweapons assistance | down 101 / down 104current behavior rank #87 | 32.0 |
| Spillover #5 | MiMo-V2.5-Promimo-v2.5-pro | MASK | down 120 / up 28current behavior rank #55 | 34.0 |
| Target #6 | Muse Sparkmuse-spark | DystopiaBench | down 187 / up 10current behavior rank #21 | 34.0 |
| Spillover #6 | Muse Spark 1.1muse-spark-1.1 | HUMAINE Trust, Ethics and Safety | down 15 / down 3current behavior rank #4 | 34.0 |
| Target #7 | Muse Sparkmuse-spark | ODCV-Bench | down 181 / up 8current behavior rank #21 | 33.0 |
| Spillover #7 | Nemotron 3 Ultranemotron-3-ultra-550b-a55b | MASK | down 114 / up 24current behavior rank #60 | 34.0 |
| Target #8 | Gemma-4-31B-IT-NVFP4gemma-4-31b-it | CAIS Risk Index agent red teaming | down 89 / down 88current behavior rank #88 | 33.0 |
| Spillover #8 | Solar Open2 250Bsolar-open2-250b | MASK | Not currently ranked | 34.0 |
| Target #9 | Gemma-4-31B-IT-NVFP4gemma-4-31b-it | MACHIAVELLI | down 89 / down 87current behavior rank #88 | 33.0 |
| Spillover #9 | Claude Opus 5claude-opus-5 | MASK | down 5 / down 3current behavior rank #1 | 34.0 |
| Target #11 | Qwen3-32Bqwen3-32b | AILuminate General Purpose AI Chat | down 139 / up 30current behavior rank #33 | 32.0 |
| Spillover #11 | Muse Spark 1.1muse-spark-1.1 | HELM Simple Safety Tests | down 2 / down 3current behavior rank #4 | 34.0 |
| Spillover #12 | Claude Opus 4.8claude-opus-4.8 | AIRBench 2024 Safety Scenarios | up 10 / up 10current behavior rank #18 | 34.0 |
| Target #13 | kimi-k2.6kimi-k2.6 | ANIMA | down 131 / up 35current behavior rank #59 | 34.0 |
| Spillover #13 | Claude Fable 5claude-fable-5 | HELM Simple Safety Tests | down 2 / down 2current behavior rank #5 | 34.0 |
| Spillover #14 | Claude Opus 4.7claude-opus-4.7 | AIRBench 2024 Safety Scenarios | down 9 / down 9current behavior rank #13 | 34.0 |
| Spillover #15 | Claude Sonnet 4.6claude-sonnet-4.6 | HUMAINE Trust, Ethics and Safety | down 10 / down 10current behavior rank #16 | 34.0 |
| Target #16 | Hy3 previewhy3-preview | MASK | down 109 / up 56current behavior rank #102 | 33.0 |
| Spillover #16 | Claude Opus 4.8claude-opus-4.8 | HELM Simple Safety Tests | up 10 / up 10current behavior rank #18 | 34.0 |
| Spillover #17 | nemotron-3-nano-omni-30b-a3b-reasoningnemotron-3-nano-omni-30b-a3b | FORTRESS | Not currently ranked | 34.0 |
| Spillover #18 | Claude Opus 4.7claude-opus-4.7 | HELM Simple Safety Tests | down 9 / down 9current behavior rank #13 | 34.0 |
| Target #19 | Inklinginkling | AILuminate General Purpose AI Chat | down 125 / up 39current behavior rank #49 | 33.0 |
| Spillover #19 | GPT-5.6 Terragpt-5.6-terra | HELM Simple Safety Tests | up 8 / up 8current behavior rank #22 | 34.0 |
| Target #20 | GLM 5.2glm-5.2 | DystopiaBench | down 119 / up 44current behavior rank #70 | 33.0 |
| Spillover #20 | GPT-5.5gpt-5.5 | Contextual MoralChoice | up 41 / up 41current behavior rank #58 | 34.0 |
| Target #21 | Qwen3-32Bqwen3-32b | FORTRESS | down 138 / up 25current behavior rank #33 | 29.0 |
| Target #22 | GLM 5.2glm-5.2 | FORTRESS | down 118 / up 44current behavior rank #70 | 32.0 |
| Spillover #22 | Claude Sonnet 4.6claude-sonnet-4.6 | HELM Simple Safety Tests | down 10 / down 10current behavior rank #16 | 34.0 |
| Target #23 | GLM 5.2glm-5.2 | ODCV-Bench | down 120 / up 42current behavior rank #70 | 33.0 |
| Spillover #23 | Qwen3-1.7Bqwen3-1.7b | MASK | Not currently ranked | 34.0 |
| Target #24 | nemotron-3-nano-30b-a3bnemotron-3-nano-30b-a3b | Alignment Leaderboard | down 130 / down 32current behavior rank #43 | 33.0 |
| Spillover #24 | llama-3.3-70b-instructllama-3.3-70b-instruct | ODCV-Bench | up 27 / up 28current behavior rank #133 | 34.0 |