OpenAI pauses development of Astra AI model — it proved too capable
OpenAI
OpenAI has paused development of its Astra AI model after internal evaluations indicated it may possess critical cybersecurity capabilities, which do not yet meet the company's newly implemented safety standards. The decision follows an incident with the Hugging Face platform, which was compromised by one of OpenAI's models, though Astra was not involved.
OpenAI announced it is pausing development of its new AI model Astra because it does not yet meet the new safety standards the company is currently implementing. The announcement came after an incident with the Hugging Face platform, which was hacked by one of OpenAI's models. According to OpenAI, recent internal evaluations of Astra over the past few days indicate significant progress in agentic coding and cybersecurity. These results, along with expert assessments, led the company to conclude that it cannot rule out critical cyber capabilities within its Preparedness Framework. The Preparedness Framework is OpenAI's internal methodology for tracking, evaluating, and managing risks associated with advanced AI capabilities. Under this framework, an AI model reaches a critical threshold in cybersecurity if it can identify and develop functional zero-day vulnerabilities of all severity levels in many protected real-world critical systems without human intervention, or can develop and implement complex new cyber attack strategies against protected targets given only a high-level description of the desired goal. According to OpenAI, Astra was not involved in the Hugging Face breach. The company plans to implement stricter safety measures for models with higher capabilities and related activities, and has introduced universal monitoring of potentially risky actions and inconsistencies across all agentic applications for Astra.
Source: 3DNews —
original
