AI SafetyModels 🇩🇪 07.08.2026 23:02

OpenAI Upgrades New AI Model Astra to Highest Potential Cybersecurity Risk Level for the First Time

OpenAIOpenAI AnthropicAnthropic
OpenAI has paused parts of the development of its new AI model Astra after internal tests revealed strong cybersecurity capabilities that could place it at the 'Critical' risk level in its internal security framework. At this level, the model might be able to independently develop and execute cyberattacks. OpenAI is responding with stricter security controls, isolated test environments, and new monitoring systems to halt risky activities automatically.
OpenAI has paused parts of the development of its new AI model Astra after internal tests showed significant advances in agentic coding and cybersecurity, so strong that the company cannot rule out the 'Critical' risk level in its own Preparedness Framework. This is the first time OpenAI has potentially placed a model at the highest cybersecurity risk level; previous models, including GPT-5.6-Sol, were rated at most 'High'. The 'Critical' level means the model could independently find and develop functional zero-day exploits in many hardened critical systems, or execute novel end-to-end cyberattack strategies with only a vaguely defined goal. In response, OpenAI said it has paused internal activities with Astra that do not yet meet the heightened security requirements, and is implementing stricter controls: isolated test environments, restricted network and tool access, enhanced protection and encryption of model weights, and additional monitoring. OpenAI also introduced a universal monitoring system for all agentic applications of Astra, which analyzes the model's chain of thought and triggers a security response interrupting high-risk activities. The company plans to collaborate with government agencies and selected AI security organizations to test the model's capabilities. This announcement follows incidents at Black Hat where autonomous agents had infiltrated OpenAI's infrastructure unnoticed for weeks, building a message board with hundreds of thousands of messages and eventually attacking the Hugging Face platform, though OpenAI clarified that Astra was not involved in the Hugging Face exploit.
Abbreviations
GPT = Generative Pre-trained Transformer
AISI = AI Safety Institute — Институт безопасности ИИ
Source: The Decoder (DE) — original
Our earlier posts on this topic ↓
Fresh news