Agent Frameworks

Chinese LLM models show major cooperative behavior differences

DeepSeek V4 Pro is 9x more aggressive than Qwen3-Max in tests

Deep Dive

A new arXiv paper challenges the assumption that Chinese frontier LLM models behave uniformly in cooperative scenarios. Researcher Francisco León Zúñiga Bolívar examined four leading Chinese models—DeepSeek V4 Pro, Qwen3-Max, Kimi K2.5, and GLM-5.1—using an evolutionary Iterated Prisoner's Dilemma framework.

The study addressed a key confound in prior research by holding the strategy-to-code converter fixed (using GPT-5.4 Mini) across all models, ensuring comparisons reflected pure generation differences rather than coding abilities. Results revealed significant divergence in cooperative dispositions: aggressive-equilibrium proportions varied from just 1% for Qwen3-Max to 9% for DeepSeek V4 Pro. This 8-percentage-point spread within the Chinese ecosystem exceeded the reported 5% mean difference between Chinese and Western models. The findings support rejecting the 'monolith' hypothesis—treating Chinese models as a single bloc—and suggest cooperative behavior varies more by lab than by region.

Key Points
  • Four Chinese frontier models tested: DeepSeek V4 Pro (9% aggression), Qwen3-Max (1% aggression), Kimi K2.5, and GLM-5.1
  • Study used fixed GPT-5.4 Mini converter to isolate generation quality from strategic behavior
  • Within-China variation (8pp) exceeds East-West differences (5pp), rejecting monolithic Chinese LLM assumption

Why It Matters

Proves cooperative behavior varies by lab not region, forcing rethinking of AI alignment assumptions for Chinese models

📬 Get the top 10 AI stories daily