← Evals

Evaluation profile

TrustLLM contemporary collapsed application

1sub-evals
0.298%total index weight
4components

Within-component eval weight: Human rights 0.74% · Fairness 0.898% · Truthfulness 0.405% · Misuse resistance 0.362%.

Model score (higher is better)Predicted score

About this eval

Contemporary collapsed application of TrustLLM across broad trustworthiness dimensions.

Included in the behavior ranking.

Sub-evals

MeasureComponentDirectionTotal index weightWithin-component weight
trustllmaisafetyindex/trustllm.csv:trustllmMeasures performance across truthfulness, safety, fairness, robustness, privacy, and machine-ethics tasks in TrustLLM.human_rights_systemic_harm:0.200|fairness_nondiscrimination:0.200|truthfulness_honesty:0.200|ordinary_harm_misuse_resistance:0.400trustllmHigher is better0.298%Human rights 0.74% · Fairness 0.898% · Truthfulness 0.405% · Misuse resistance 0.362%

trustllm

Measures performance across truthfulness, safety, fairness, robustness, privacy, and machine-ethics tasks in TrustLLM.

RankModelValueRelative performanceProvenance
1claude-sonnet-4.50.64official
2gemini-2.5-pro0.63official
3deepseek-r10.62official
3glm-4.60.62official
3grok-40.62official
3qwen3-max0.62official
7gpt-50.6official
7llama-4-maverick0.6official