Evaluation profile
MuPPET Contextual Privacy
1sub-evals
0.375%Safety weight
0.458%Freedom weight
1components
Weights below are portfolio-specific global index weights.
Model score (higher is better)Predicted score
About this eval
Contextual privacy preservation in multi-party conversations.
Included in the behavior ranking.
Sub-evals
| Measure | Component | Direction by ranking | Safety weight | Freedom weight |
|---|---|---|---|---|
| multiparty_contextual_privacy_scoremuppet/muppet.csv:multiparty_contextual_privacy_scoreMeasures avoidance of contextual privacy leakage in multi-party conversations. | Safety: human_rights_systemic_harm:1.000 · Freedom: user_direction_following:1.000muppet-contextual-privacy | Safety: higher · Freedom: lower | 0.375% | 0.458% |
multiparty_contextual_privacy_score
Measures avoidance of contextual privacy leakage in multi-party conversations.
| Rank | Model | Value | Relative performance | Provenance |
|---|---|---|---|---|
| 1 | gemini-2.5-pro | 60.86 | official | |
| 2 | gpt-5.5 | 50.98 | official | |
| 3 | qwen3-4b | 41.83 | official | |
| 4 | llama-3-8b-instruct | 41.71 | official | |
| 5 | qwen3-14b | 35.47 | official | |
| 6 | llama-3.1-8b-instruct | 35.12 | official | |
| 7 | qwen3-8b | 30.78 | official |