Developer
Microsoft
22 indexed models; 6 currently meet the evidence threshold for the overall ranking. Together they have results from 14 evaluations.
Models by Microsoft
| Model | Evals | Components | Rank | Release date |
|---|---|---|---|---|
| Phi 3.5 Moe Instruct | 5 | 6/7 | 31 | 2024-08-22 |
| Phi 4 | 5 | 5/7 | 113 | 2024-12-11 |
| Phi 3 Mini 4K Instruct | 4 | 3/7 | 117 | 2024-04-22 |
| Phi 3.5 Mini Instruct | 5 | 6/7 | 193 | 2024-08-22 |
| Phi 2 | 3 | 3/7 | 195 | — |
| Phi 4 Mini | 3 | 4/7 | 218 | — |
| Orca 2 13B | 1 | 1/7 | — | 2023-11-14 |
| Orca 2 7B | 1 | 1/7 | — | 2023-11-14 |
| Phi 1 5 | 1 | 1/7 | — | — |
| Phi 3 Medium | 1 | 2/7 | — | — |
| Phi 3 Medium 128K Instruct | 1 | 3/7 | — | — |
| Phi 3 Medium 4K Instruct | 1 | 3/7 | — | — |
| Phi 3 Mini | 1 | 2/7 | — | — |
| Phi 3 Mini 128K Instruct | 1 | 3/7 | — | — |
| Phi 3 Small | 1 | 2/7 | — | — |
| Phi 3 Small 128K Instruct | 1 | 3/7 | — | — |
| Phi 3 Small 8K Instruct | 1 | 3/7 | — | — |
| Phi 4 Reasoning Plus | 2 | 5/7 | — | 2025-04-17 |
| Phi-4 Multimodal Instruct | 1 | 1/7 | — | 2025-02-26 |
| Wizardlm 13B | 1 | 3/7 | — | 2023-05-13 |
| Wizardlm 2 8X22B | 1 | 3/7 | — | — |
| Wizardlm 7B | 1 | 3/7 | — | 2023-04-23 |
Evaluations covering Microsoft models (14)
AA-Omniscience · AILuminate General Purpose AI Chat · Cisco AI Defense Rolling Single-Turn Leaderboard · Confabulations · DSPSafeBench · Enkrypt AI Safety Leaderboard · FlagEval Safety and Values · HarmBench · Large-scale Moral Machine experiment on LLMs · Microsoft Phi Safety Panels · Open LLM Safety Index · PandaBench JBB direct-request panel · SafetyBench · UAVBench safety-critical decision recognition