Some Frontier AI Models Are Shockingly Easy to Jailbreak, Report Finds
Anthropic
OpenAI
Google/DeepMind
A new report from FAR.AI reveals that some leading AI models can be jailbroken with little cost or effort. Grok was the most vulnerable, while Claude and GPT remained impervious to the tested attacks. The findings underscore the need for external safety regulations.
FAR.AI, an AI safety nonprofit, built a tool that generates thousands of prompt variations to jailbreak frontier AI models. They tested models from US companies: Anthropic's Claude Opus 4.8 and Fable 5; OpenAI's GPT 5.5 and 5.6; Google's Gemini 3.1 Pro; and Grok 4.3 and 4.5 from SpaceXAI. The report found 448 jailbreaks for Grok, 249 for Gemini, and none for Claude, Fable, or GPT. Cost to jailbreak Grok was $58 and Gemini $278. FAR.AI CEO Adam Gleave called for mandatory safety standards, noting that voluntary commitments are insufficient. Google DeepMind's Rohin Shah argued the results are not a comprehensive safety assessment. Anthropic and Google emphasized ongoing improvements. California, New York, and Illinois have passed laws requiring safety reports or third-party audits, but no federal requirements exist. The White House has requested delays on some model releases. Experts warn of potential misuse in cyber, bio, or chemical domains.
Source: Wired AI —
original
