Language Models Cannot Trigger Scientific Revolutions, but World Models Might
Google DeepMind
OpenAI
Google/DeepMind
Sakana AI
A position paper by Google DeepMind's Tom Zahavy argues that large language models (LLMs) lack the cognitive mechanism needed for scientific discovery, specifically 'manipulative abduction' — the creative leap to new axioms. Zahavy suggests world models that allow active intervention in simulations might be a path forward.
In a position paper titled 'LLMs can't jump,' Tom Zahavy of Google DeepMind argues that language models cannot trigger scientific revolutions because they lack the key cognitive mechanism for creating genuinely new ideas. He uses Einstein's model of discovery as a cycle: from sensory experience via an intuitive 'jump' to axioms, then logical derivation to testable conclusions. Zahavy distinguishes three types of reasoning: deduction, induction, and abduction, and claims that LLMs handle induction and deduction but not the 'manipulative abduction' — inventing a cause for which there is no linguistic template yet. He illustrates with Einstein's work, noting that there was no data crisis at the time; Newton's physics was highly accurate except for a tiny anomaly in Mercury's orbit, and a purely optimizing AI might have invented a hypothetical planet rather than rethinking space and time. Zahavy argues that Einstein's 'happiest thought' — the freely falling observer — came from embodied simulation, a bodily feeling, not formula manipulation. He compares this to Archimedes' eureka moment in the bathtub. Language models lack this sensory grounding; they are like the 'Chinese Room' — moving symbols without access to embodied experience. He notes that systems like Sakana's AI Scientist and DeepMind's AlphaEvolve can automate science but either recombine existing concepts or require a clear error signal, which Einstein never had. As a potential path, Zahavy suggests physically consistent world models: video generators like Veo only predict the next frame, but action-controllable world models like Genie allow agents to intervene and run counterfactual experiments, providing the feedback needed to invent new axioms. Zahavy is cautious, saying the paper 'suggests' but does not prove that this jump is the critical bottleneck.
- Сокращения
- LLM = Large Language Model — большая языковая модель
Source: The Decoder (DE) —
original
