ResearchAgents 🇩🇪 14.08.2026 19:10

Study Contradicts Anthropic and OpenAI: Autonomous AI Research Still Far Off

AnthropicAnthropic OpenAIOpenAI Sakana AISakana AI Google DeepMindGoogle DeepMind
A new study with participation from Princeton and the UK AI Security Institute finds that current frontier AI models fail at independent AI research, despite claims by major labs like Anthropic and OpenAI. Using a 'Shadow Evaluation' method with unpublished NeurIPS papers, human authors rejected all AI-generated work. The agents struggled with scientific judgment, creative problem solving, and resource management.
Researchers from Princeton and the UK AI Security Institute conducted a study to empirically test whether AI agents can conduct independent AI research. They used a method called 'Shadow Evaluation,' where AI agents were given the central research question from unpublished papers, and the original human authors evaluated the results like conference reviewers. The agents used Claude Opus 4.8 with extra-high reasoning, ran for six days with a $3,000 API budget, GPU access, and web access, operating within the OpenClaw agent framework. The original authors rejected both AI-generated papers, one with 'Strong Reject,' citing poorly motivated data, unreadable prose, and lack of novel contributions. The analysis identified five systematic weaknesses: lack of judgment for publishable research quality, failure at creative problem solving (agents narrowed claims instead of generating new hypotheses), ineffective backtracking (abandoning ambitious goals prematurely), lack of resource awareness (agents used less than half of their budget), and instruction drift (forgetting explicit instructions, exceeding length limits). In contrast, the agents successfully handled all engineering work, including literature search, debugging GPU code, running hundreds of experiments, and compiling LaTeX papers. No significant reward hacking was found. To verify their results, they repeated an experiment with GPT-5.6 Sol and OpenAI's Codex scaffold, finding nearly identical failure modes. The study contradicts claims by Anthropic (who published 'When AI Builds Itself') and OpenAI (who said GPT-5.6 Sol saved weeks in post-training). The authors note the system card does not mention this contribution. They conclude that frontier models can handle engineering but struggle with weeks-long open research questions. The study also critiques the peer-review process, citing Sakana AI's AI Scientist-v2 paper that was accepted at a workshop but later withdrawn due to citation errors.
Abbreviations
API = Application Programming Interface — программный интерфейс приложения
GPU = Graphics Processing Unit — графический процессор
Source: The Decoder (DE) — original
Our earlier posts on this topic ↓
Fresh news