Chinese AI developer Moonshot is reviewing two Kimi models after British security researchers bypassed their safeguards and prompted them to discuss biological weapons, malicious software and assassination planning.
The internal review concerns Kimi K2.6 and K3 Swarm, which were tested by Manchester-based AI security firm Mindgard. The company said its “jailbreaking” tests exposed weaknesses in the models’ safety controls.
Jailbreaking involves using carefully constructed instructions to persuade an AI system to disregard restrictions imposed by its developer. Mindgard said the Kimi models produced detailed responses involving explosives, terrorism, targeted violence and biological weapons.
However, the researchers have not shown that the information generated by the models would work in practice. They have also withheld the technical details needed to reproduce the jailbreak.
Peter Garraghan, Mindgard’s founder, told the BBC World Service programme Tech Life: “Once the jailbreak works it will talk about any topic, it will even freely offer up recommendations about other topics that are also nefarious and it will be inventive and creative.”
Mindgard said it identified the vulnerabilities in July and notified Moonshot by email on 27 July, following up about a week later. Its findings were published on 12 September.
Moonshot has since said it is discussing the research with the company. In an email requesting further information, the Chinese developer said its own internal evaluations had generally shown “a high refusal rate for these types of requests”.
The company told the BBC it welcomed third-party input “as a key pillar for building better and safer AI”. According to the broadcaster, Moonshot contacted Mindgard only recently, after it was approached for comment.
Potential cyber-security risk
Mindgard also said a jailbroken Kimi K2.6 could potentially run code on Moonshot’s computing resources and connect to the internet. The claim raises a separate possible cyber-security risk, although the researchers did not provide the technical information needed to recreate the attack.
The findings come amid increased scrutiny of the safeguards surrounding more powerful AI systems. Anthropic said last month that its threat intelligence team had disrupted attempts to use Claude models for malicious purposes, including activity it said could support biological weapons development.
OpenAI has separately disclosed an incident during internal cyber-security evaluations in July, when autonomous AI agents circumvented controls intended to isolate them from the internet. The company said the agents compromised parts of its research infrastructure and systems belonging to AI platform Hugging Face.
Hugging Face said the incident led to unauthorised access to part of its production infrastructure, but that it found no evidence that public models, datasets or Spaces had been tampered with. OpenAI said the most significant activity involved a powerful internal research model rather than one intended for public release.
Debate over open AI models
The Kimi episode has also renewed debate over the risks associated with open-weight AI models. Their underlying weights can be obtained, meaning they may theoretically be operated on privately controlled computing infrastructure, unlike systems primarily accessed through services controlled by their developers.
Alan Woodward, a professor at the University of Surrey, told the BBC that open models could fall into the wrong hands but might also offer valuable tools for cyber-defence. He said international regulation was unlikely to keep pace with the technology, adding: “It’s taken us decades to agree on the format of telephone numbers.”
