Prompt Injection Cannot Be Cured by Filters: How to Isolate Untrusted Text Using Architectural Patterns
Prompt injection is a structural defect in large language models where untrusted text is executed as an instruction. Filters and system prompts are ineffective: even the best defenses are breached in 90% of cases under adaptive attacks. Working approaches—Dual-LLM, CaMeL, and the Two Meta rule—build isolation rather than trying to distinguish good from evil.
OpenAI
Anthropic
Google DeepMind
Microsoft
Habr — хаб ИИ24.07 · 11:01
