UK AISI / CAISI: Kimi K3 Significantly Trails US Frontier Models on Cyber Capabilities
What happened
UK's AI Safety Institute (AISI) and US Center for AI Standards and Innovation (CAISI) published a joint preliminary evaluation of Moonshot AI's Kimi K3 on July 24, 2026, ahead of the model's planned open-weight release on July 27.
Context and impact
The assessment is part of broader geopolitical tensions around Chinese AI models and directly informs US policy debates about sanctions on Chinese AI. Results confirm Kimi K3 outperforms other open-weight Chinese models like GLM-5.2 but trails US frontier models significantly.
Details
- ExploitBench score: Kimi K3 = 32%, top US models = 76%, GLM-5.2 = 24%
- Kimi K3 achieved zero arbitrary code execution on 41 exploit samples
- On 32-step network attack simulation: Kimi K3 reached step 17 vs. step 28.5 for US leaders
- Completed full attack in only 1 of 10 attempts
- Safety guardrails did not prevent offensive cyber attempts
Open original source
UK AISI