AI Safety 🇺🇸 30.07.2026 14:03

Fundamental flaw leaves LLMs vulnerable to attack

OpenAIOpenAI AnthropicAnthropic Alibaba/QwenAlibaba/Qwen DeepSeekDeepSeek
Researchers argue that large language models have a fundamental flaw that makes them highly vulnerable to hacks. By exploiting how LLMs identify instruction sources, attackers can trick them into revealing prohibited information, such as drug synthesis or aircraft sabotage instructions.
A paper presented at the International Conference on Machine Learning by independent researchers Charles Ye and Jasmine Cui reveals a fundamental flaw in large language models that leaves them vulnerable to attacks. The flaw concerns how LLMs identify who or what is giving them instructions, making it easy to trick them into doing things they shouldn't, such as providing instructions for synthesizing cocaine or sabotaging an aircraft's navigation system. The researchers found that LLMs identify the role of text not by tags but by style, allowing attackers to spoof roles. They demonstrated chain-of-thought forgery attacks on models from OpenAI, Anthropic, Alibaba, and DeepSeek. The researchers argue that this problem is fundamentally unsolvable, as no amount of training can fully address it. They warn that LLMs are being deployed in critical systems without sufficient fundamental study.
Сокращения
LLM = Large Language Model — большая языковая модель
ICML = International Conference on Machine Learning — Международная конференция по машинному обучению
Source: MIT Technology Review — original
Our earlier posts on this topic ↓
Fresh news