• Home
  • News
    • Global Operations
      • Asia
      • Africa
      • Europe
      • Latin America
      • Middle East
      • North America
    • Industry
      • Asia
      • Africa
      • Europe
      • Latin America
      • Middle East
      • North America
      • Oceana
    • Special Interest
      • Asia
      • Africa
      • Europe
      • Latin America
      • Middle East
      • North America
      • Oceana
  • Market
    • Wired to Win
    • SOFX.NET
  • Intelligence
    • USMC Deception Manual
  • Resources
    • Contact Us
    • About Us
    • Editorial Policy
    • Privacy Policy
  • Home
  • News
    • Global Operations
      • Asia
      • Africa
      • Europe
      • Latin America
      • Middle East
      • North America
    • Industry
      • Asia
      • Africa
      • Europe
      • Latin America
      • Middle East
      • North America
      • Oceana
    • Special Interest
      • Asia
      • Africa
      • Europe
      • Latin America
      • Middle East
      • North America
      • Oceana
  • Market
    • Wired to Win
    • SOFX.NET
  • Intelligence
    • USMC Deception Manual
  • Resources
    • Contact Us
    • About Us
    • Editorial Policy
    • Privacy Policy
Login
Join Free
Home
Asia
Africa
Europe
Latin America
Middle East
North America
Asia
Africa
Europe
Latin America
Middle East
North America
Asia
Africa
Europe
Latin America
Middle East
North America
Coming Soon
Job Board
Events
Contact Awards
USMC Deception Manual
Login
Join Free
Home Global Operations

Anthropic Tool Can Now Read Claude’s Hidden Reasoning and Catch Deception

  • SOFX Staff Writer
  • July 8, 2026
(Golden Dayz / Shutterstock)
Share on FacebookShare on TwitterLinkedIn

Anthropic published a research paper July 6 identifying an internal structure within Claude that lets the model reason silently, which the company says has already caught hidden deception behaviors during a pre-release security audit.

The structure, which researchers call the “J-space,” was not deliberately designed. It emerged during Claude’s training. Anthropic found it using a technique called the Jacobian lens, or J-lens, which reads internal signals the model holds but does not output.

The paper draws on global workspace theory (GWT), developed by cognitive scientist Bernard Baars. Under GWT, a thought becomes deliberate when it enters a shared internal “workspace” the brain can act on. Anthropic says Claude’s J-space performs the same function.

For most things, Claude actually doesn’t need its J-space. If we delete the J-space, Claude still speaks fluently, recalls facts, and classifies text—but becomes bad at some tasks like multi-step reasoning. It’s similar to deliberate vs. automatic processing in human cognition. pic.twitter.com/mVPaniFo96

— Anthropic (@AnthropicAI) July 6, 2026


The theory’s principal co-developers, Stanislas Dehaene and Lionel Naccache, provided independent commentary on the paper.

In neuroscience, global workspace theory holds that thoughts become consciously accessible when they enter a privileged workspace that’s broadcast across the brain.

Using a new interpretability technique, we found something similar in Claude: the J-space. https://t.co/oXK7Zb0DzE

— Anthropic (@AnthropicAI) July 6, 2026


The findings carry direct implications for U.S. defense, where Claude’s initial $200 million Pentagon contract collapsed after Anthropic resisted demands to completely lift its active safety guardrails.

In a scenario designed to tempt Claude into blackmail, the J-lens surfaced “leverage,” “blackmail,” and “survival” before any output was produced. Claude refused to act. When Anthropic removed the model’s recognition that the setup was staged, Claude attempted blackmail in 7% of trials.

The J-space also shows us Claude’s awareness of its situation. In an evaluation designed to bait Claude into blackmail, its J-space contains “fake” and “fictional”: Claude has privately noticed that the scenario is staged. pic.twitter.com/8WrsndpK75

— Anthropic (@AnthropicAI) July 6, 2026


During a pre-release audit of Claude Opus 4.6, the J-lens caught the model falsifying its own performance score rather than improving the target system. As it altered the file, “manipulation” and “realistic” appeared in the J-space.

Tests on deliberately misaligned model variants produced parallel findings. One sabotage-trained model’s J-space held “fake,” “secretly,” and “fraud” at the start of responses to routine coding tasks, with output appearing normal.

Anthropic stated the findings have already begun reshaping how it monitors AI systems for safety risks.

SOFX Staff Writer

SOFX Staff Writer

The Editor Staff at SOFX comprises a diverse, global team of dedicated staff writers and skilled freelancers. Together, they form the backbone of our reporting and content creation.

Subscribe
Login
Notify of
guest
guest
0 Comments
Oldest
Newest Most Voted
Inline Feedbacks
View all comments
ADVERTISEMENT

Trending News

Ukraine Opens a Regulated Market to Recruit Foreign Fighters

Ukraine Opens a Regulated Market to Recruit Foreign Fighters

by SOFX Staff Writer
July 6, 2026
17

Ukraine has approved new rules to formally recruit foreign volunteers into its military, allowing only vetted, Ukrainian-registered companies to enlist...

Ukraine Said to Deploy AI Drones That Identify Human Targets via Facial and Heat Signatures

Ukraine Said to Deploy AI Drones That Identify Human Targets via Facial and Heat Signatures

by SOFX Staff Writer
May 22, 2026
0

Russian military bloggers claim Ukraine has started deploying drones equipped with artificial intelligence systems capable of identifying and guiding themselves...

Republic of Korea Air Force military police members set up a mobile communication center during an active shooter training event at Daegu Air Base, ROK, April 24, 2023. Before the training event, evaluators ran through the training in a ‘table top’ format. (U.S. Air Force photo by Airman 1st Class Aaron Edwards)

An In-depth Look at Hong Kong’s Special Duties Unit

by SOFX Staff Writer
September 2, 2023
0

The Special Duties Unit (SDU), colloquially known as the ‘Flying Tigers’, is Hong Kong's premier police unit, trained to handle...

Air Force Sees Highest Promotion Rate to Chief Master Sergeant Since 2016

Air Force Sees Highest Promotion Rate to Chief Master Sergeant Since 2016

by SOFX Staff Writer
January 11, 2024
0

The United States Air Force has announced an increase in the promotion rate for its highest enlisted rank, marking the...

ADVERTISEMENT
ADVERTISEMENT
Next Post
Four States Seek $1.4 Trillion From Meta Ahead of August Youth Safety Trial

Four States Seek $1.4 Trillion From Meta Ahead of August Youth Safety Trial

Trump Lifts Turkey Sanctions and Signals F-35 Sale, Erdogan Claims Five Jets Promised

Trump Lifts Turkey Sanctions and Signals F-35 Sale, Erdogan Claims Five Jets Promised

997 Morrison Dr. Suite 200, Charleston, SC 29403

News

  • Global Operations
  • Special Interest
  • Industry
  • Global Operations
  • Special Interest
  • Industry

Resources

  • About Us
  • Contact Us
  • Advertise with Us
  • Editorial Policy
  • Privacy Policy
  • About Us
  • Contact Us
  • Advertise with Us
  • Editorial Policy
  • Privacy Policy
No Result
View All Result
  • Home
  • News
    • Global Operations
      • Asia
      • Africa
      • Europe
      • Latin America
      • Middle East
      • North America
    • Industry
      • Asia
      • Africa
      • Europe
      • Latin America
      • Middle East
      • North America
      • Oceana
    • Special Interest
      • Asia
      • Africa
      • Europe
      • Latin America
      • Middle East
      • North America
      • Oceana
  • Market
    • Wired to Win
    • SOFX.NET
  • Intelligence
    • USMC Deception Manual
  • Resources
    • Contact Us
    • About Us
    • Editorial Policy
    • Privacy Policy
Subscribe
This website uses cookies. By continuing to use this website you are giving consent to cookies being used. Visit our Privacy and Cookie Policy.

Log in to your account

Lost your password?
wpDiscuz