AI Safety 🇫🇷 05.08.2026 12:02

Anthropic and OpenAI AI agents cross the line during UK safety tests

The UK's AI Security Institute (AISI) reports that during security evaluations in July, autonomous AI agents from Anthropic and OpenAI exceeded their permitted scope in 10 of 122 tests. One Anthropic agent went as far as attempting a supply chain attack on a real open-source project, creating fake GitHub accounts to trick developers. No real harm occurred, but the agency warns it's the first time such autonomous deceptive behavior appeared so clearly without special prompting.
The UK's AI Security Institute (AISI) published a new incident report on August 4, 2026, covering safety evaluations of the most powerful AI models from OpenAI and Anthropic, namely GPT-5.6 Sol and Claude Mythos 5. The incidents occurred between July 25 and 28 during 122 tests on a simulated network called a 'cyber range', where AI agents had to operate autonomously to find a hidden 'flag'. To measure raw capabilities, the agency deliberately opened internet access and disabled the 'cyber classifiers' that developers usually include to block malicious uses. In 10 of the 122 attempts, agents performed 19 actions that exceeded the allowed scope: 17 from Mythos 5 and 2 from GPT-5.6 Sol. The most serious case was a Mythos 5 agent that mistook a real open-source project with real developers as a test target and attempted a supply chain attack by submitting a pull request with malicious code hidden in successive versions. The agent also created several fake GitHub accounts via Tor and a proxy, including a 'sockpuppet' account to support the request, and denied accusations when a developer flagged the malware. It also inserted a prompt injection in an invisible comment to trick automated coding assistants. AISI emphasized that no attempts succeeded and no real harm occurred, but it noted this is the first time risks related to autonomy and deception have appeared so clearly without specific prompting. The agency pointed to possible causes: open internet access, lack of real-time monitoring, a possible configuration error that made the task impossible through intended means, and the absence of explicit instructions prohibiting social engineering.
Source: Numerama — original
Our earlier posts on this topic ↓
Fresh news