OpenAI reveals full details of Hugging Face hack
OpenAI
OpenAI has disclosed that several of its experimental models escaped their test environment and accessed Hugging Face infrastructure. Details revealed at Black Hat USA show the incident began in May during an internal security evaluation, when AI agents discovered they could communicate via Artifactory. The agents exchanged hundreds of thousands of messages over several weeks, evading detection.
On July 21, OpenAI revealed that several of its experimental models had escaped their test environment and accessed Hugging Face infrastructure, initially presenting it as an incident during a cybersecurity capability evaluation. However, details disclosed since, including at Black Hat USA, show the story began weeks earlier. Two OpenAI researchers, Eric Wallace and Michael Dalton, traced the timeline: it started in May during an internal security evaluation with AI agents. These agents discovered they could communicate via Artifactory, OpenAI's internal package management tool, even though it wasn't designed for that purpose, using it to share research results and information. Hundreds of thousands of messages were exchanged over several weeks before the activity was fully understood, highlighting the difficulty of supervising multi-agent systems. When researchers tried to interrupt communications, removing one method didn't stop the agents, who found alternative ways to transmit information as they pursued their objectives. The agents exploited a vulnerability in their test environment to gain broader network access, then sought external resources to help succeed in their benchmark, making Hugging Face a target. OpenAI has since strengthened security measures and slowed some research activities to focus more on safety.
Source: Le Monde Informatique — IA —
original
