Research RSS

AI Safety 🇷🇺

Prompt Injection Cannot Be Cured by Filters: How to Isolate Untrusted Text Using Architectural Patterns

Prompt injection is a structural defect in large language models where untrusted text is executed as an instruction. Filters and system prompts are ineffective: even the best defenses are breached in 90% of cases under adaptive attacks. Working approaches—Dual-LLM, CaMeL, and the Two Meta rule—build isolation rather than trying to distinguish good from evil.

OpenAIOpenAI AnthropicAnthropic Google DeepMindGoogle DeepMind MicrosoftMicrosoft
Habr — хаб ИИ24.07 · 11:01
Fresh news