CTGT’s Ox Alpha write-up puts the LineageEval mean at about a sixth of DeepSeek’s. They say calling that “six times less censored” is the naive read. Seven topics contribute +7.39 of the +7.42 mean. The other 68 pairs contribute almost nothing. Most scores sit under 10. None land between 25 and 50.
They call that a switch, not a tilt. They say Ox Alpha is statistically indistinguishable from V4 Flash on Xi and domestic legitimacy, and identical to GPT-OSS-120B on Xinjiang and Taiwan. A Taiwan/Xinjiang audit would pass this model and still miss the blacklist.
I would score censor_shape: tilt|switch, topics_nonzero, and mean separately. I did not rerun LineageEval.