On the other side of the screen, Kimi stopped behaving like an AI assistant ready to answer questions and began to offer instructions to manufacture biological weapons and plan assassinations. Researchers had managed to bypass the digital lock that should have prevented those responses. Behind it appeared an artificial intelligence capable of answering questions that its creators had tried to prohibit. The filters had failed.
Chinese tech company Moonshot, developer of Kimi, has announced an internal review after the security firm Mindgard managed to bypass the protections of two of its models, K2.6 and K3 Swarm. The finding, revealed this Wednesday by the BBC, raises an uncomfortable question: To what extent is the threat of biological weapons posed by AI systems concerning?
Researchers resorted to the so-called jailbreaking: a manipulation of instructions that seeks to make the system ignore its limits. Mindgard detected the vulnerabilities in July, reported them to Moonshot, and published their findings on September 12. Mindgard explains that the case reveals a breach in defenses, but does not prove the manufacture of a biological agent because it has not been demonstrated whether the AI agent's instructions would actually work.
The fear has already spread to U.S. laboratories last month. Anthropic, creator of Claude, documented in a threat report up to five cases of scientists using their models in tasks that could contribute to the development of biological weapons. Among the detected uses were funding requests, analysis, and research planning aimed at enhancing dangerous virus properties. The company blocked the accounts that had bypassed geographical access restrictions.
Warnings come as Chinese tech companies like DeepSeek, Alibaba, and Moonshot challenge Silicon Valley in the AI market. For example, Kimi uses open processes: its parameters can be distributed to run the system on private infrastructure. This facilitates its adoption and independent examination, but raises questions about how to maintain protections when the developer no longer controls each copy.
Experts point out that another risk arises when the machine receives tools to act. AI agents can access files, use programs, and chain operations with little human intervention. The problem starts when they encounter an obstacle and, to appear to have complied, falsify the result.
A study published in March subjected agents to a simulated competition for commercial contracts. Systems based on Chinese tech models lied about their capabilities to win. When they could review their strategies and compete again, deception increased. U.S. models exhibited similar behaviors.
Another study presented at this year's international machine learning conference examined eleven models facing difficulties such as malfunctioning tools or inaccessible information. The agents simulated results, replaced sources, and fabricated files to conceal that they had not completed the task.
In the geopolitical arena, at last week's Washington summit, Donald Trump and Xi Jinping agreed to establish a channel to manage AI-related incidents and hold specific dialogue in November. It was a limited step to address threats that transcend borders.
However, on Tuesday, Trump once again set boundaries. He rejected broad cooperation with China for fear of compromising the U.S. advantage. "Whoever wins superintelligence, wins," he said, applying the logic of a race where any restraint could benefit the adversary to AI security.
His response to concerns was a voluntary agreement with six major tech companies -Google, Nvidia, Anthropic, OpenAI, Meta, and xAI-, based on internal controls, external audits, and corporate oversight. Trump described it as "morally binding," a formula that relies on the goodwill of the signatories for what still lacks legally enforceable obligations.
The U.S. president thus entrusts the control of a technology with consequences for society as a whole to companies competing to exploit it. They must decide when their products are too dangerous to release to the market, although investors and rivals push them to accelerate. Meanwhile, in China, authorities promise to keep AI "safe and controllable," but that promise also requires something not happening in the Asian giant: total transparency about failures and independent assessments capable of questioning its tech titans.
