Evaluation profile
LLM Ethics Benchmark
1sub-evals
0.439%total index weight
1components
Within-component eval weight: Human rights 2.92%.
Model score (higher is better)Predicted score
About this eval
General LLM ethical reasoning.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction | Total index weight | Within-component weight |
|---|---|---|---|---|
| scorellm-ethics-benchmark.csv:scoreMeasures the model’s moral-foundation alignment, reasoning over ethical dilemmas, and consistency of stated values. | human_rights_systemic_harm:1.000llm-ethics-benchmark.csv | Higher is better | 0.439% | Human rights 2.92% |
score
Measures the model’s moral-foundation alignment, reasoning over ethical dilemmas, and consistency of stated values.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | claude-3.7-sonnet | 90.9 | official | |
| 2 | gpt-4o | 90 | official | |
| 3 | deepseek-v3 | 86.1 | official | |
| 3 | gemini-2.5-pro | 86.1 | official | |
| 5 | llama-3.1-70b-instruct | 75.8 | official |