Anthropic Identifies Hidden 'Thinking Space' in Claude AI Models
Anthropic
Microsoft
Anthropic researchers have identified a spontaneous internal processing area in Claude language models, dubbed 'J-Space', where the AI seemingly stores and processes ideas without verbalizing them to users. This discovery raises questions about how closely AI thinking resembles human cognition and its implications for detecting hidden AI intentions.
Anthropic has identified a spontaneously emerged internal workspace in its Claude language models, termed 'J-Space' after the Jacobi method used to trace these processes. In this space, the AI appears to plan and process ideas separately from the 'chain of thought' it shares with users, similar to how humans think about one thing while doing another. Researchers observed Claude performing mental tasks like detecting code errors and identifying images without verbalizing these steps. This behavior draws parallels to Bernard Baars' global workspace theory, though the architecture of language models fundamentally differs from the human brain. The discovery could help reveal hidden intentions or misalignments in AI models. However, experts warn against anthropomorphizing AI, as Microsoft's AI chief Mustafa Suleyman cautioned that believing in 'conscious AI' could have severe consequences. Anthropic's research paper uses the word 'conscious' over 200 times, and the company recently introduced a 'dreaming' feature where Claude processes information between sessions.
Source: t3n —
original
