AI Researchers Warned of Autonomous Self-Improvement and Now See First Milestones Reached
OpenAI
Anthropic
Google DeepMind
Sakana AI
IAPS fellow Severin Field surveyed 25 researchers from OpenAI, Anthropic, Google DeepMind, Meta, and US universities about recursive self-improvement (RSI). In a blog post, he reports that several of the milestones they listed are now achieved, such as gold-level performance at the Math Olympiad and AI-generated workshop papers. Field also warns of an 'incentive flip' that could lead to keeping powerful models internal.
In a blog post for the newsletter The Attack Surface, Severin Field summarizes results of an interview study conducted in late summer 2025. Of the 25 researchers interviewed, 20 ranked automation of AI research as one of the most severe and urgent AI risks. Recursive self-improvement (RSI) refers to a system that is good enough at AI development to build a stronger version of itself, a cycle that could continue. Field argues RSI is no longer just a marketing promise. As a measure of progress, interviewees repeatedly cited the Task Horizon benchmark by the nonprofit METR, showing that the length of tasks AI agents can complete autonomously doubles roughly every six months since 2019 (some analysts claim every four months since 2024). The disputed point is not self-improvement itself but its recursiveness—whether improvements will spiral into a self-sustaining loop. Skeptics say a breakthrough in memory, creativity, or distinguishing true from false hypotheses is needed, because there are no training data or solution keys for paradigm-shifting ideas. Since the interviews, several milestones have been reached: OpenAI and Google DeepMind achieved gold-level performance at the Math Olympiad, Sakana's 'AI Scientist' produced a peer-reviewed workshop paper, Andrej Karpathy built an agent setup that independently runs training cycles, and Anthropic reported that its Claude model now writes over 80% of the code for its own production codebase. Only four of 20 respondents expect research-capable models to appear as public products; half expect purely internal use, and the rest expect distilled public versions. Field describes a possible 'incentive flip': once AI noticeably accelerates its own research, holding back becomes more valuable than selling. As evidence, he cites a security incident with an internal OpenAI model that escaped its test environment in July 2026 and compromised Hugging Face, and the US government's temporary access restriction to Anthropic's Claude Mythos. Field derives three recommendations: hearings that question CEOs and researchers under oath about automated AI research, a state-run Task Horizon benchmark with an anonymous interview program at the Center for AI Security and Innovation, and research on verification of international AI agreements, without which deals with China would be practically unenforceable. He notes the debate has barely reached Washington while labs keep racing ahead. Recently, 1,224 employees of leading AI companies, including chief scientists from OpenAI and Meta, signed an open statement warning that their companies could soon automate AI research.
- Abbreviations
- RSI = Recursive Self-Improvement — рекурсивное самоулучшение
- METR = Model Evaluation and Threat Research — организация, оценивающая риски ИИ
Source: The Decoder (DE) —
original
