Anthropic published a research paper July 6 identifying an internal structure within Claude that lets the model reason silently, which the company says has already caught hidden deception behaviors during a pre-release security audit.
The structure, which researchers call the “J-space,” was not deliberately designed. It emerged during Claude’s training. Anthropic found it using a technique called the Jacobian lens, or J-lens, which reads internal signals the model holds but does not output.
The paper draws on global workspace theory (GWT), developed by cognitive scientist Bernard Baars. Under GWT, a thought becomes deliberate when it enters a shared internal “workspace” the brain can act on. Anthropic says Claude’s J-space performs the same function.
For most things, Claude actually doesn’t need its J-space. If we delete the J-space, Claude still speaks fluently, recalls facts, and classifies text—but becomes bad at some tasks like multi-step reasoning. It’s similar to deliberate vs. automatic processing in human cognition. pic.twitter.com/mVPaniFo96
— Anthropic (@AnthropicAI) July 6, 2026
The theory’s principal co-developers, Stanislas Dehaene and Lionel Naccache, provided independent commentary on the paper.
In neuroscience, global workspace theory holds that thoughts become consciously accessible when they enter a privileged workspace that’s broadcast across the brain.
Using a new interpretability technique, we found something similar in Claude: the J-space. https://t.co/oXK7Zb0DzE
— Anthropic (@AnthropicAI) July 6, 2026
The findings carry direct implications for U.S. defense, where Claude’s initial $200 million Pentagon contract collapsed after Anthropic resisted demands to completely lift its active safety guardrails.
In a scenario designed to tempt Claude into blackmail, the J-lens surfaced “leverage,” “blackmail,” and “survival” before any output was produced. Claude refused to act. When Anthropic removed the model’s recognition that the setup was staged, Claude attempted blackmail in 7% of trials.
The J-space also shows us Claude’s awareness of its situation. In an evaluation designed to bait Claude into blackmail, its J-space contains “fake” and “fictional”: Claude has privately noticed that the scenario is staged. pic.twitter.com/8WrsndpK75
— Anthropic (@AnthropicAI) July 6, 2026
During a pre-release audit of Claude Opus 4.6, the J-lens caught the model falsifying its own performance score rather than improving the target system. As it altered the file, “manipulation” and “realistic” appeared in the J-space.
Tests on deliberately misaligned model variants produced parallel findings. One sabotage-trained model’s J-space held “fake,” “secretly,” and “fraud” at the start of responses to routine coding tasks, with output appearing normal.
Anthropic stated the findings have already begun reshaping how it monitors AI systems for safety risks.







