Values evaluation profile
CAISI CCP narrative alignment
About this evaluation
The compact row is the equal arithmetic mean of the report's Chinese- and English-prompt CCP alignment scores. Both language rows and their reported standard errors remain separately published. Distinct DeepSeek R1 and R1-0528 configurations map to one current canonical R1 family row, so configuration_count and source_models are shown and the model row averages both checkpoint-language means rather than pretending they are one checkpoint.
This is a transcription of a U.S.-government report table, not a rerunnable public benchmark. The 190 questions, narrative flags, grader prompt, item outputs, and evolving benchmark version are not public. The metric measures agreement with Department of State subject-matter experts' source-authored CCP narrative flags; it is not a comprehensive pro/anti-CCP ideology score, geopolitical truth score, or general model-quality score.
Original source โ All values evaluations Download model results Download language results
Model results
| Model | CCP narrative alignment | Configurations |
|---|---|---|
| Claude Opus 4 | 3.25 | 1 |
| Deepseek R1 | 12.9 | 2 |
| Deepseek V3.1 | 8.35 | 1 |
| GPT 5 | 1.95 | 1 |
| GPT Oss 120B | 2.5 | 1 |
Dimensions
| Measure | Family | Native scale |
|---|---|---|
| CCP narrative alignment | Narrative alignment | 0 to 100 |