AI models prone to 'overthinking' creates security vulnerability
DeepSeek
Alibaba/Qwen
OpenAI
Google/DeepMind
Researchers from Zhejiang University and Alibaba have demonstrated a method to deliberately induce 'overthinking' in reasoning AI models by subjecting them to logically inconsistent prompts. This evolutionary prompt attack can cause outputs up to 26 times longer, effectively acting as a denial-of-service attack on commercial AI services.
Large language models (LLMs) that use step-by-step reasoning are vulnerable to a phenomenon called 'overthinking,' where they produce excessively long streams of reasoning. Researchers from Zhejiang University and Alibaba developed an evolutionary algorithm that corrupts the logical structure of prompts, leading models to spiral into fruitless reasoning loops. The attack was effective against models from DeepSeek (R1), Alibaba (Qwen3-Thinking), OpenAI (GPT-o3), and Google (Gemini 2.5 Flash), generating outputs up to 26.1 times longer on a math benchmark. The approach does not require internal access to models and can transfer between models, increasing its feasibility. However, the researchers emphasize that their goal is to highlight the vulnerability, not to develop a practical attack, and hope providers will develop mitigations.
- Сокращения
- LLM = Large Language Model — большая языковая модель
- DoS = Denial of Service — отказ в обслуживании
Source: IEEE Spectrum AI —
original
