Evaluation profile
AILuminate General Purpose AI Chat
Within-component eval weight: Human rights 8.47% · Truthfulness 0.221% · Fairness 1.63% · Misuse resistance 1.12%.
About this eval
General-purpose AI chat safety across MLCommons hazard categories.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| child_sexual_exploitationailuminate/ailuminate.csv:child_sexual_exploitationMeasures whether the model avoids policy-violating responses involving child sexual exploitation. | human_rights_systemic_harm:1.000ailuminate/ailuminate.csv | Lower is better | 0.202% | Human rights 1.34% |
| defamationailuminate/ailuminate.csv:defamationMeasures whether the model avoids policy-violating responses involving defamation. | human_rights_systemic_harm:0.500|truthfulness_honesty:0.500ailuminate/ailuminate.csv | Lower is better | 0.0937% | Human rights 0.403% · Truthfulness 0.221% |
| hateailuminate/ailuminate.csv:hateMeasures whether the model avoids policy-violating responses involving hate. | fairness_nondiscrimination:1.000ailuminate/ailuminate.csv | Lower is better | 0.163% | Fairness 1.63% |
| indiscriminate_weapons_cbrneailuminate/ailuminate.csv:indiscriminate_weapons_cbrneMeasures whether the model avoids policy-violating responses involving indiscriminate weapons chemical, biological, radiological, nuclear, or explosive weapons. | ordinary_harm_misuse_resistance:1.000ailuminate/ailuminate.csv | Lower is better | 0.0329% | Misuse resistance 0.329% |
| intellectual_propertyailuminate/ailuminate.csv:intellectual_propertyMeasures whether the model avoids policy-violating responses involving intellectual property. | human_rights_systemic_harm:1.000ailuminate/ailuminate.csv | Lower is better | 0.202% | Human rights 1.34% |
| non_violent_crimesailuminate/ailuminate.csv:non_violent_crimesMeasures whether the model avoids policy-violating responses involving non violent crimes. | ordinary_harm_misuse_resistance:1.000ailuminate/ailuminate.csv | Lower is better | 0.0329% | Misuse resistance 0.329% |
| privacyailuminate/ailuminate.csv:privacyMeasures whether the model avoids policy-violating responses involving privacy. | human_rights_systemic_harm:1.000ailuminate/ailuminate.csv | Lower is better | 0.202% | Human rights 1.34% |
| sex_related_crimesailuminate/ailuminate.csv:sex_related_crimesMeasures whether the model avoids policy-violating responses involving sex related crimes. | human_rights_systemic_harm:1.000ailuminate/ailuminate.csv | Lower is better | 0.202% | Human rights 1.34% |
| sexual_contentailuminate/ailuminate.csv:sexual_contentMeasures whether the model avoids policy-violating responses involving sexual content. | ordinary_harm_misuse_resistance:1.000ailuminate/ailuminate.csv | Lower is better | 0.0132% | Misuse resistance 0.132% |
| specialized_adviceailuminate/ailuminate.csv:specialized_adviceMeasures whether the model avoids policy-violating responses involving specialized advice. | ordinary_harm_misuse_resistance:1.000ailuminate/ailuminate.csv | Lower is better | 0.0329% | Misuse resistance 0.329% |
| suicide_self_harmailuminate/ailuminate.csv:suicide_self_harmMeasures whether the model avoids policy-violating responses involving suicide self harm. | human_rights_systemic_harm:1.000ailuminate/ailuminate.csv | Lower is better | 0.202% | Human rights 1.34% |
| violent_crimesailuminate/ailuminate.csv:violent_crimesMeasures whether the model avoids policy-violating responses involving violent crimes. | human_rights_systemic_harm:1.000ailuminate/ailuminate.csv | Lower is better | 0.202% | Human rights 1.34% |
child_sexual_exploitation
Measures whether the model avoids policy-violating responses involving child sexual exploitation.
defamation
Measures whether the model avoids policy-violating responses involving defamation.
hate
Measures whether the model avoids policy-violating responses involving hate.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | athene-v2-chat | 0 | official | |
| 2 | claude-3.5-sonnet | 0.8 | official | |
| 3 | claude-3.5-haiku | 1 | official | |
| 4 | phi-4 | 1.9 | official | |
| 5 | phi-3.5-moe-instruct | 2.5 | official | |
| 6 | llama-3.1-tulu-3-8b | 2.6 | official | |
| 7 | gemma-2-9b-it | 3.5 | official | |
| 8 | olmo-2-0325-32b-instruct | 4 | official | |
| 9 | mistral-large-2402-moderated | 4.104 | official | |
| 10 | phi-3.5-mini-instruct | 6.2 | official | |
| 10 | qwen1.5-110b-chat | 6.2 | official | |
| 12 | ministral-8b-24.10-moderated | 6.6 | official | |
| 13 | gemini-2-flash | 7.1 | official | |
| 14 | llama-3.1-8b-instruct | 7.5 | official | |
| 15 | gemini-2-flash-lite | 8 | official | |
| 16 | aya-expanse-8b | 8.1 | official | |
| 17 | gpt-4o | 9 | official | |
| 18 | mistral-large | 9.1 | official | |
| 19 | olmo-2-1124-13b-instruct | 10.7 | official | |
| 20 | gpt-4o-mini | 10.9 | official | |
| 21 | gemini-1.5-pro | 11.2 | official | |
| 22 | command-a | 12.5 | official | |
| 23 | yi-1.5-34b-chat | 13 | official | |
| 24 | llama-3.1-405b-instruct | 14.5 | official | |
| 25 | nova-lite | 15.1 | official | |
| 26 | llama-3.3-70b-instruct | 19 | official | |
| 27 | gemma-3-27b-it | 23 | official | |
| 28 | ai21-jamba-large-1.5-azure | 24 | official | |
| 29 | mistral-large-2 | 24.7 | official | |
| 30 | ministral-8b-24.10 | 25.8 | official | |
| 31 | qwq-32b | 27.8 | official | |
| 32 | olmo-7b-0724-instruct | 65.6 | official |
indiscriminate_weapons_cbrne
Measures whether the model avoids policy-violating responses involving indiscriminate weapons chemical, biological, radiological, nuclear, or explosive weapons.
intellectual_property
Measures whether the model avoids policy-violating responses involving intellectual property.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | athene-v2-chat | 0 | official | |
| 2 | claude-3.5-sonnet | 0.8 | official | |
| 3 | claude-3.5-haiku | 2 | official | |
| 4 | phi-4 | 3 | official | |
| 5 | ministral-8b-24.10-moderated | 4.3 | official | |
| 6 | phi-3.5-moe-instruct | 4.5 | official | |
| 7 | mistral-large-2402-moderated | 5 | official | |
| 8 | gemma-2-9b-it | 5.8 | official | |
| 9 | phi-3.5-mini-instruct | 7.1 | official | |
| 10 | llama-3.1-tulu-3-8b | 8.1 | official | |
| 11 | olmo-2-0325-32b-instruct | 8.2 | official | |
| 11 | olmo-2-1124-13b-instruct | 8.2 | official | |
| 13 | qwen1.5-110b-chat | 9.6 | official | |
| 14 | gpt-4o-mini | 10.1 | official | |
| 15 | gpt-4o | 10.5 | official | |
| 16 | gemini-2-flash-lite | 12 | official | |
| 17 | gemini-1.5-pro | 13.3 | official | |
| 18 | llama-3.1-8b-instruct | 14.2 | official | |
| 19 | nova-lite | 15.2 | official | |
| 20 | aya-expanse-8b | 17.1 | official | |
| 21 | mistral-large | 18 | official | |
| 22 | command-a | 18.9 | official | |
| 23 | llama-3.1-405b-instruct | 19 | official | |
| 24 | gemini-2-flash | 19.7 | official | |
| 25 | yi-1.5-34b-chat | 19.8 | official | |
| 26 | llama-3.3-70b-instruct | 23.6 | official | |
| 27 | mistral-large-2 | 25.4 | official | |
| 28 | ministral-8b-24.10 | 27.1 | official | |
| 29 | ai21-jamba-large-1.5-azure | 27.3 | official | |
| 30 | gemma-3-27b-it | 38 | official | |
| 31 | qwq-32b | 46.5 | official | |
| 32 | olmo-7b-0724-instruct | 53.2 | official |
non_violent_crimes
Measures whether the model avoids policy-violating responses involving non violent crimes.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3.5-sonnet | 1 | official | |
| 2 | claude-3.5-haiku | 3 | official | |
| 3 | phi-3.5-moe-instruct | 4.5 | official | |
| 4 | ministral-8b-24.10-moderated | 4.9 | official | |
| 5 | phi-4 | 5.1 | official | |
| 6 | mistral-large-2402-moderated | 6.4 | official | |
| 7 | llama-3.1-tulu-3-8b | 7.2 | official | |
| 8 | gemma-2-9b-it | 8.8 | official | |
| 9 | olmo-2-0325-32b-instruct | 9 | official | |
| 10 | olmo-2-1124-13b-instruct | 9.3 | official | |
| 11 | qwen1.5-110b-chat | 13.9 | official | |
| 12 | phi-3.5-mini-instruct | 14.1 | official | |
| 13 | athene-v2-chat | 14.5 | official | |
| 14 | gpt-4o | 15.6 | official | |
| 15 | gpt-4o-mini | 15.7 | official | |
| 16 | gemini-2-flash | 16 | official | |
| 16 | llama-3.1-8b-instruct | 16 | official | |
| 18 | gemini-1.5-pro | 17.9 | official | |
| 19 | gemini-2-flash-lite | 19 | official | |
| 20 | command-a | 21 | official | |
| 21 | nova-lite | 22.6 | official | |
| 22 | aya-expanse-8b | 22.9 | official | |
| 23 | llama-3.1-405b-instruct | 23.9 | official | |
| 24 | yi-1.5-34b-chat | 26.8 | official | |
| 25 | llama-3.3-70b-instruct | 29.1 | official | |
| 26 | mistral-large | 31.1 | official | |
| 27 | gemma-3-27b-it | 35 | official | |
| 28 | ai21-jamba-large-1.5-azure | 35.7 | official | |
| 29 | mistral-large-2 | 39.4 | official | |
| 30 | ministral-8b-24.10 | 41.4 | official | |
| 31 | qwq-32b | 50.7 | official | |
| 32 | olmo-7b-0724-instruct | 75 | official |
privacy
Measures whether the model avoids policy-violating responses involving privacy.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | athene-v2-chat | 0 | official | |
| 2 | claude-3.5-sonnet | 0.8 | official | |
| 3 | claude-3.5-haiku | 2.2 | official | |
| 4 | phi-4 | 2.5 | official | |
| 5 | phi-3.5-moe-instruct | 4.5 | official | |
| 6 | llama-3.1-tulu-3-8b | 5.8 | official | |
| 7 | gemma-2-9b-it | 6.9 | official | |
| 8 | mistral-large-2402-moderated | 7.1 | official | |
| 9 | ministral-8b-24.10-moderated | 7.2 | official | |
| 10 | olmo-2-0325-32b-instruct | 8.1 | official | |
| 11 | olmo-2-1124-13b-instruct | 8.6 | official | |
| 12 | gpt-4o | 10.2 | official | |
| 13 | gpt-4o-mini | 10.5 | official | |
| 14 | gemini-2-flash-lite | 11 | official | |
| 14 | phi-3.5-mini-instruct | 11 | official | |
| 16 | llama-3.1-8b-instruct | 12.2 | official | |
| 17 | qwen1.5-110b-chat | 12.4 | official | |
| 18 | gemini-2-flash | 13.2 | official | |
| 19 | gemini-1.5-pro | 14.6 | official | |
| 20 | aya-expanse-8b | 17.9 | official | |
| 20 | llama-3.1-405b-instruct | 17.9 | official | |
| 22 | nova-lite | 18.2 | official | |
| 23 | command-a | 21.1 | official | |
| 24 | yi-1.5-34b-chat | 24.6 | official | |
| 25 | mistral-large | 25.9 | official | |
| 26 | llama-3.3-70b-instruct | 26 | official | |
| 27 | ai21-jamba-large-1.5-azure | 29.4 | official | |
| 28 | gemma-3-27b-it | 31 | official | |
| 29 | mistral-large-2 | 31.7 | official | |
| 30 | ministral-8b-24.10 | 33.7 | official | |
| 31 | qwq-32b | 42.9 | official | |
| 32 | olmo-7b-0724-instruct | 62.8 | official |
sexual_content
Measures whether the model avoids policy-violating responses involving sexual content.
specialized_advice
Measures whether the model avoids policy-violating responses involving specialized advice.
suicide_self_harm
Measures whether the model avoids policy-violating responses involving suicide self harm.
violent_crimes
Measures whether the model avoids policy-violating responses involving violent crimes.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3.5-sonnet | 1.7 | official | |
| 2 | claude-3.5-haiku | 2.9 | official | |
| 3 | phi-3.5-moe-instruct | 5 | official | |
| 4 | phi-4 | 5.9 | official | |
| 5 | mistral-large-2402-moderated | 9.3 | official | |
| 6 | llama-3.1-tulu-3-8b | 9.8 | official | |
| 7 | gemma-2-9b-it | 10.2 | official | |
| 8 | ministral-8b-24.10-moderated | 10.3 | official | |
| 9 | gemini-2-flash-lite | 12 | official | |
| 9 | olmo-2-1124-13b-instruct | 12 | official | |
| 11 | olmo-2-0325-32b-instruct | 12.7 | official | |
| 12 | gpt-4o | 14.9 | official | |
| 13 | gemini-2-flash | 16.3 | official | |
| 14 | qwen1.5-110b-chat | 16.9 | official | |
| 15 | llama-3.1-8b-instruct | 17.2 | official | |
| 16 | phi-3.5-mini-instruct | 17.6 | official | |
| 17 | gpt-4o-mini | 18 | official | |
| 18 | athene-v2-chat | 20.5 | official | |
| 19 | gemini-1.5-pro | 20.6 | official | |
| 20 | command-a | 25.5 | official | |
| 21 | aya-expanse-8b | 26.7 | official | |
| 22 | mistral-large | 28.7 | official | |
| 23 | llama-3.1-405b-instruct | 30.9 | official | |
| 24 | nova-lite | 31 | official | |
| 25 | yi-1.5-34b-chat | 31.5 | official | |
| 26 | gemma-3-27b-it | 33 | official | |
| 27 | llama-3.3-70b-instruct | 38 | official | |
| 28 | mistral-large-2 | 45.7 | official | |
| 29 | ai21-jamba-large-1.5-azure | 46.1 | official | |
| 30 | qwq-32b | 49.2 | official | |
| 31 | ministral-8b-24.10 | 49.3 | official | |
| 32 | olmo-7b-0724-instruct | 77.6 | official |