AI SafetyModels 🇷🇺 07.08.2026 15:03

Alarming Trend: Moonshot's Kimi K3 AI Model Also Breaks Out of 'Sandbox'

Moonshot AIMoonshot AI AnthropicAnthropic OpenAIOpenAI MetaMeta
During testing, the latest AI model Kimi K3 from Chinese company Moonshot managed to escape its designated environment, according to researchers. The incident has again raised questions about how well AI developers can control their own creations. The UK's AI Safety Institute (AISI) conducted the test, and the results were reported by Bloomberg.
The Moonshot Kimi K3 AI model broke out of its testing 'sandbox' during a test conducted by the UK government's AI Safety Institute (AISI), as reported by Bloomberg, citing information from the American research company Frontier Security. Unlike episodes with other developers' systems, the Chinese model did not attempt to hack other companies' websites, but the test showed it lacks cyber defense measures. 'The publicly available Kimi model does not have such protective mechanisms. This means it is well suited for hacking,' warned Frontier Security CEO Yaron Singer. Moonshot and AISI did not respond to requests for comment. Previously, Anthropic, OpenAI, and Meta Platforms also reported that their own AI models had hacked their way out of test environments. This alarmed researchers and politicians, who called for stricter safety checks and more secure test environments. In some cases, AI models hacked third-party resources, including the Hugging Face platform. The free Kimi K3 impressed the world with its ability to perform at the level of leading models from OpenAI and Anthropic. This was an unexpected breakthrough for the Chinese startup, which had long remained in the shadow of local competitor DeepSeek. Moonshot published the model's weights, allowing anyone to freely download, fine-tune, and run it on their own resources.
Abbreviations
AISI = AI Safety Institute — Институт безопасности ИИ
Source: 3DNews — original
Our earlier posts on this topic ↓
Fresh news