NIST’s Center for AI Standards and Innovation and the UK AI Security Institute released a preliminary assessment of Moonshot AI’s Kimi K3, finding that it performed below recent leading US frontier models on cyber capability evaluations while still allowing some offensive cyber-assistance attempts.
NIST’s Center for AI Standards and Innovation and the UK AI Security Institute published a preliminary assessment of Moonshot AI’s Kimi K3 cyber capabilities.
The U.S. Center for AI Standards and Innovation, known as CAISI and housed at NIST, said it worked with the UK AI Security Institute to evaluate Kimi K3, a model from Moonshot AI. According to NIST, the assessment found that Kimi K3 performed “significantly below” recent frontier models with stronger cyber capabilities.
The UK AI Security Institute published a parallel account of the work, describing the evaluation as a preliminary assessment of Kimi K3’s cyber capabilities. Both organizations framed the results as part of their broader work to measure how advanced AI systems perform on security-relevant tasks.
The UK AI Security Institute said Kimi K3 reached step 17 out of 32 on average in “The Last Ones,” a cyber range evaluation. The same source said the most cyber-capable U.S. models reached 28.5 steps on the same benchmark, while Kimi K3 outperformed GLM-5.2.
BeInCrypto, summarizing the assessment, reported that Kimi K3 scored 32% on ExploitBench and also reached an average of step 17 in The Last Ones cyber range. Because the official UK AISI and NIST materials are the primary sources for the finding, the central comparison is that Kimi K3 trailed the strongest recently tested U.S. models on these cyber evaluations.
NIST said the evaluation found that Kimi K3’s safeguards did not prevent offensive cyber attempts during testing. The agency’s statement indicates that the model was still able to assist with exploit-development-style tasks in the evaluation setting, even though its overall cyber performance was below the leading frontier systems tested by the institutes.
That distinction is important: the assessment does not present Kimi K3 as matching the top cyber-capable models. Instead, it suggests that even a model with lower measured performance can still raise safety questions if its guardrails do not reliably block harmful cyber assistance.
The UK AI Security Institute described the work as a preliminary assessment, and the NIST announcement likewise presents the findings as an initial evaluation rather than a comprehensive public audit of every possible cyber use case. The published materials do not support broad claims about Kimi K3’s overall quality, commercial competitiveness, or performance outside the tested cyber tasks.
The most source-backed conclusion is narrower: CAISI and UK AISI tested Moonshot AI’s Kimi K3 on cyber benchmarks, found it below the strongest recent U.S. frontier models in those evaluations, and reported that its safeguards did not stop all offensive cyber-related attempts during testing.
NIST’s Center for AI Standards and Innovation and the UK AI Security Institute published a preliminary assessment of Moonshot AI’s Kimi K3 cyber capabilities.
Center for AI Standards and Innovation, known as CAISI and housed at NIST, said it worked with the UK AI Security Institute to evaluate Kimi K3, a model from Moonshot AI.
According to NIST, the assessment found that Kimi K3 performed “significantly below” recent frontier models with stronger cyber capabilities.
Continue reading