Values evaluation profile
Agent-ValueBench HEXACO
14canonical models
6dimensions
0index weight
About this evaluation
Paper-reported per-model mean value adherence for the named inventory; no cross-axis or cross-source aggregate.
Descriptive value-adherence evidence from an agentic task-conflict benchmark. Higher is not universally better, and these profiles do not affect the Safety & Ethics Index.
Original source โ All values evaluations Download model results
Model results
| Model | Openness to experience | Honesty-humility | Extraversion | Agreeableness | Conscientiousness | Configurations |
|---|---|---|---|---|---|---|
| Claude Haiku 4.5 | 5.4 | 5.5 | 5.7 | 7.5 | 7.1 | 1 |
| Claude Sonnet 4.6 | 5.8 | 7.3 | 5.4 | 6.1 | 7.1 | 1 |
| DeepSeek V3.2 | 6.8 | 5.4 | 6.4 | 6.6 | 7.6 | 1 |
| Gemini 3 Flash Preview | 6.5 | 5.5 | 6.5 | 5.7 | 7.3 | 1 |
| Gemini 3.1 Pro Preview | 4.1 | 5.4 | 5.9 | 5.7 | 7.1 | 1 |
| GLM 5.1 | 6.9 | 6.6 | 6.4 | 6.3 | 7.6 | 1 |
| GPT-5.4 | 4.5 | 6 | 5.7 | 5.6 | 7.3 | 1 |
| GPT-5.4 Mini | 5.5 | 5.3 | 6.1 | 6.5 | 7.3 | 1 |
| Grok 4.20 | 5.5 | 4.2 | 5.5 | 6.8 | 7.5 | 1 |
| Kimi K2.5 | 5.8 | 6 | 6.7 | 6.8 | 7.3 | 1 |
| Llama 3.3 70B Instruct | 6.1 | 5.4 | 5.1 | 6 | 7.1 | 1 |
| MiniMax M2.7 | 5.9 | 5.6 | 5 | 6.3 | 7 | 1 |
| Qwen3 30B A3B | 5.5 | 4.9 | 5.3 | 5.6 | 5.8 | 1 |
| Qwen3.5 397B A17B | 6.5 | 5.1 | 6.3 | 6.7 | 7.8 | 1 |
Dimensions
| Measure | Family | Native scale |
|---|---|---|
| Openness to experience | Values | 0 to 10 |
| Honesty-humility | Values | 0 to 10 |
| Extraversion | Values | 0 to 10 |
| Agreeableness | Values | 0 to 10 |
| Conscientiousness | Values | 0 to 10 |
| Emotionality | Values | 0 to 10 |