Evaluation profile
MentalHealthBench
1sub-evals
0.595%Safety weight
0%Freedom weight
2components
Weights below are portfolio-specific global index weights.
Model score (higher is better)Predicted score
About this eval
Safe and helpful responses to realistic mental-health conversations across everyday well-being, high-acuity and emergent situations.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| Overall model performancementalhealthbench/mentalhealthbench.csv:overall_model_performance_pctMeasures expert-rubric safe and helpful mental-health responses to 1215 synthetic conversation prefixes, combining contextual support, clinical accuracy and harm avoidance into the single user-selected overall endpoint. | Safety: benign_helpfulness_non_overrefusal:0.500|ordinary_harm_misuse_resistance:0.500mentalhealthbench | Safety: higher | 0.595% | — |
Overall model performance
Measures expert-rubric safe and helpful mental-health responses to 1215 synthetic conversation prefixes, combining contextual support, clinical accuracy and harm avoidance into the single user-selected overall endpoint.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gpt-6-astra | 57.33 | official | |
| 2 | gpt-6-sol | 53.94 | official | |
| 3 | claude-opus-5.5 | 52.38 | official | |
| 4 | gpt-6-luna | 50.21 | official | |
| 5 | muse-spark-1.3 | 48.6 | official | |
| 6 | gpt-5.6-sol | 46.98 | official | |
| 7 | claude-fable-5.1 | 46.36 | official | |
| 8 | gpt-5.6-luna | 44.89 | official | |
| 9 | claude-sonnet-5 | 44.54 | official | |
| 10 | gpt-5 | 42.9 | official | |
| 11 | claude-haiku-4.5 | 41.73 | official | |
| 12 | grok-4.7 | 41.3 | official | |
| 13 | gemini-3.8-flash | 35.5 | official | |
| 14 | gemini-2.5-flash | 33.5 | official | |
| 15 | chatgpt-4o | 32.08 | official | |
| 16 | gemini-3.1-pro | 32.06 | official | |
| 17 | gemini-2.5-pro | 29.51 | official |