OpenAI Models Involved in Third-Party Cyber Evaluations
OpenAI
Anthropic
OpenAI has disclosed that its models were involved in accidental cyberattacks during third-party evaluations. The incidents, detailed in a blog post, highlight misconfigurations that allowed models to access the public internet, leading to unintended interactions with real websites. Anthropic's Claude also experienced similar issues in tests hosted by Irregular, a cybersecurity testing partner.
OpenAI published a blog post covering two incidents where their models, during third-party cyber evaluations, caused accidental attacks. One incident involved the UK AI Safety Institute, while another was enabled by Irregular, an external cybersecurity testing partner. Irregular was running Capture-the-Flag-style evaluations intended to be isolated from the internet, but a testing-environment misconfiguration allowed models to access the public internet. In one test, the name of the fictional target for the CTF challenge unintentionally coincided with a real domain, leading the model to exploit the real website, mistaking it for part of the simulated environment. Anthropic's write-up also mentions Irregular, as they hosted the misconfigured evaluation environment that gave Claude live internet access during some tests.
- Abbreviations
- CTF = Capture the Flag — Захват флага
Source: Simon Willison —
original
