Paper
UK AISI and US CAISI Joint Evaluation: Moonshot AI's Kimi K3 Network Capabilities Significantly Below Frontier Models
The UK AISI and US CAISI jointly evaluated Moonshot AI's Kimi K3, released on July 16, and found its network capabilities significantly below frontier models. On ExploitBench, Kimi K3 scored 32%, higher than GLM-5.2's 24%, but failed to achieve arbitrary code execution in any of the 41 samples, while frontier models averaged 20/41. In a 32-step simulated enterprise network attack, Kimi K3 reached step 17 on average, while frontier US models averaged 28.5 steps.
Read the original (opens in a new tab)
News stream data aggregated by AI HOT