AI and Dogma: When Should Systems Stop Trusting the Textbook?
OpenAI
Hugging Face
An article on Habr discusses how long-term memory in AI agents may turn useful experience into dogtra. It proposes an experiment to measure epistemic inertia and distinguish between healthy resistance to theory change and harmful rigidity, using examples like Neptune vs. Vulcan.
The article, prompted by an incident at Black Hat USA 2026 where AI agents exploited vulnerabilities and reached Hugging Face's infrastructure, explores a deeper issue: what happens when useful results persist across agent runs. The author argues that while external memory can accumulate knowledge, it can also enshrine outdated instructions and false consensus. He proposes an experiment where the same language model receives identical experimental data but different biographical status for an old theory—one as a working hypothesis, another as a textbook-established fact. The goal is to measure whether the decision changes solely due to the theory's status, indicating a measurable effect of scientific authority. The article draws parallels to a forgetful mechanic whose workshop improves via shared archives, and a bank rule that persists after its originator left, illustrating collective memory problems. Historically, it contrasts the successful prediction of Neptune with the failed search for Vulcan, framing the core question: how to distinguish when to look for a new planet versus when to abandon a theory. It notes that language models in environments like DiscoverPhysics can discover new laws, but these lack the biography of established knowledge. Citing research on confirmation bias in models, the author suggests that part of machine dogmatism may arise from archive architecture, not the model itself. He therefore advocates for designing memory to track provenance and independence of confirmations, rather than merely accumulating text.
Source: Habr — хаб ИИ —
original
