Enterprise & Industry

China's Kimi K3 trails US rivals in cyberattack capabilities, study finds

Kimi K3 scored just 32.2% vs 76.2% average for top US models on hacking benchmarks.

Deep Dive

A joint study by the UK Artificial Intelligence Security Institute (AISI) and the US Centre for AI Standards and Innovation (CAISI) has found that Moonshot AI's Kimi K3, considered China's most powerful large language model, performs 'significantly below' top American rivals in its ability to launch cyberattacks. The model scored just 32.2% on ExploitBench, a public benchmark for assessing AI's ability to develop cybersecurity exploits. In contrast, unnamed leading US models averaged 76.2%. Kimi K3 failed to achieve arbitrary code execution—the highest level of exploit granting full control of a target system—across all 41 ExploitBench tasks, while US models succeeded on 20 tasks.

These results challenge Washington's brewing anxiety over the rapid rise of Chinese open-source AI. While Kimi K3 outperformed domestic rival Zhipu AI's GLM-5.2 (24.4%), it still lags far behind frontier US models. The report suggests that current Chinese AI models are not yet a major cybersecurity threat, contradicting fears that open-source Chinese models could be easily weaponized for hacking.

Key Points
  • Kimi K3 scored 32.2% on ExploitBench, while top US models averaged 76.2%.
  • It failed to achieve arbitrary code execution on any of 41 tasks; US models succeeded on 20 tasks.
  • The study was conducted by UK AISI and US CAISI, challenging concerns over Chinese AI cyber capabilities.

Why It Matters

The findings ease fears about Chinese AI being a near-term cybersecurity threat, highlighting a significant gap with US frontier models.

📬 Get the top 10 AI stories daily