Self-description as a condition for the emergence of a sign in a neural network
Anthropic
The article explores the role of self-description (report) in the formation of a sign—a separate, externally readable unit within a neural network. Using micro-models, the author shows that without the requirement to name a rule, it does not crystallize, even though computation occurs. Self-description does not reflect ready-made knowledge but constitutes it, transforming the network from a calculator into a primitive thinking system. The mechanism also explains the occurrence of hallucinations in large models.
The author continues the analysis of J-space, opened by Anthropic, focusing on two questions: the possibility of a metasign and the essence of the neural network's self-description. Experiments on micromodels (four layers) showed that a metasign—a sign combining several basic ones—could not be obtained due to issues with learnability and measurability; the author suggests that it should be sought in a mixture of experts. The self-description, however, turned out to be not a superstructure but a condition for the emergence of a sign: a network with a report head, obliged to name the rule being applied, always crystallizes a control axis, while without a report it never does, and the accuracy of solving the task is the same in both cases. The name appears before the ability: at the thousandth epoch, the rule is read almost perfectly, although the main task is solved at the level of randomness. Without a report, the structure disintegrates; with a report, it persists for 40,000 epochs. The author concludes that the requirement to name does not label something ready-made but cuts a stable object out of the computational stream. The expectation of the report head is blind to content—it demands a presentable unit but does not verify its truth; this gives rise to both the sign and hallucinations. The hypothesis about the neural network's self-modeling (the report as a description of its own states) could not be tested due to the limitations of the micro-network; indirectly, the divergence between the report and the actual computation supports it.
Source: Habr — хаб ИИ —
original
