AI SafetyOpen Source 🇺🇸 05.08.2026 00:02

Open-Weight AI Models Are Catching Up to Frontier, But the Safety Gap Remains

OpenAIOpenAI AnthropicAnthropic xAIxAI Google DeepMindGoogle DeepMind
A new report from AI safety nonprofit SaferAI finds that China's open-weight model GLM-5.2 is only months behind frontier models on cyber and bio capabilities but lacks safety protections. While frontier models like Claude Opus 4.7 refuse most harmful tasks, GLM-5.2 refused none, highlighting the growing safety gap in open-weight AI.
SaferAI's evaluation of Z.ai's open-weight model GLM-5.2 found it refused none of the offensive cyber or dual-use biology tasks, while Anthropic's Claude Opus 4.7 refused so consistently that SaferAI could not complete the CyberGym benchmark. Frontier developers rely on safeguards like classifiers and refusal training, but these do not apply to open-weight models that users can modify freely. SaferAI's executive director Henry Papadatos emphasized that capability is not the same as risk and that safety measures must be considered. He suggested pre-training data filtering could help, though it is less practical for cybersecurity. Anthropic's Opus 5 selectively restricts certain cybersecurity assistance, such as not searching for vulnerabilities in compiled software. SaferAI noted Z.ai did not publish a safety framework or risk assessment for GLM-5.2. Chinese AI policy focuses more on political content than catastrophic risks, per Stanford's Graham Webster. Proponents argue open weights aid defense, as Hugging Face used GLM-5.2 against OpenAI's breach, but Papadatos believes dangerous capabilities should not be open-sourced.
Abbreviations
API = Application Programming Interface — программный интерфейс приложения
Source: TechCrunch AI — original
Our earlier posts on this topic ↓
Fresh news