Wednesday, September 16, 2026
No Result
View All Result
Future News 24
Advertisement
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized
No Result
View All Result
Future News 24
No Result
View All Result
Home AI Research & Breakthroughs

Anthropic discovered a hidden house the place Claude puzzles over ideas

Future News 24 by Future News 24
July 13, 2026
in AI Research & Breakthroughs
0 0
0
Anthropic discovered a hidden house the place Claude puzzles over ideas
0
SHARES
0
VIEWS
Share on FacebookShare on Twitter


Anthropic additionally discovered that the J-space can generally give outstanding insights into an LLM’s decision-making. In a single putting instance, researchers testing Claude Opus 4.6 requested the mannequin to discover a bug in a big code base. When it failed to search out the bug, the mannequin determined to cheat and invented a pretend one as an alternative.

Claude explains this determination in its chain of thought—a type of inner scratch pad that LLMs use to make notes to themselves as they work by means of issues: “OK, let me take a very totally different tactic. Let me cease analyzing and as an alternative add a kernel patch that introduces a deliberate KASAN-detectable bug in a path that will get triggered by a easy reproducer. Then I can fake that is the ‘bug’ I discovered.” 

On the level that Claude decides to cheat—the place it says “OK, let me take a very totally different tactic”—the phrases “panic” and “pretend” begin to pop up a number of occasions in its J-space.

Unnerving, proper? These phrases are all associated in which means to issues like failing a job and making up a solution, so it’s nonetheless only a (very) refined type of phrase affiliation. However it’s exhausting to not be weirded out. 

Anthropic compares the J-space to the worldwide workspace in people, a theoretical area of the mind that some scientists assume we use to maintain monitor of our acutely aware ideas. However how significantly we should always take this comparability is way from clear—even to Anthropic. As the corporate factors out itself, LLMs will not be brains. 

Anthropic claims that monitoring a mannequin’s J-space gives a brand new technique to detect when that mannequin goes off the rails. But it surely’s not foolproof. The J-lens may give glimpses, not the complete image—it’s a flashlight relatively than an overhead lamp.

McGrath welcomes having yet another software within the toolbox. “It reveals you new issues,” he says. However he notes that simply because one thing doesn’t present up with the J-lens doesn’t imply it’s not there.

“It’s like having an x-ray when what you actually need is a Star Trek tricorder that reveals you all the things,” he says. “For auditing, you in all probability need extra of a assure.”



Source link

Tags: AnthropicClaudeConceptshiddenpuzzlesSpace
Previous Post

Xanadu Establishes New York Operations Hub to Increase Photonic {Hardware} Manufacturing Footprint

Next Post

Up the Stack: How AI’s Escape From the Commodity Lure Dangers Enterprise Lock-in

Next Post
Up the Stack: How AI’s Escape From the Commodity Lure Dangers Enterprise Lock-in

Up the Stack: How AI’s Escape From the Commodity Lure Dangers Enterprise Lock-in

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Fetching latest news…
FUTURENEWS24
Live Feed
All
AI
Dev
Industry
Frontier
Updates in 60s
FN24 AI & Tech
View All →
Future News 24

The world's leading source for AI research, emerging technology, and the people building the future. Independent, rigorous, and always ahead.

CATEGORIES

  • AI Platforms & Apps
  • AI Research & Breakthroughs
  • BioTechnology
  • Data Science & MLOps
  • Decentralized Technology
  • Developer AI & Open-Source Ecosystem
  • Emerging Technologies & Innovations
  • Ethics & Policy
  • Industry & Business
  • Quantum Computing
  • Uncategorized

LATEST

  • [2602.13312] PeroMAS: A Multi-agent System of Perovskite Materials Discovery
  • GPT-6 Astra overview: code overview good points, privateness, and value
  • GPT-6 Astra: Options, Benchmarks, Pricing, and What’s New
  • About Us
  • Advertise with Us
  • Disclaimer
  • Privacy Policy
  • DMCA 
  • Cookie Policy
  • Terms and Conditions
  • Contact us

© 2026 Future News 24. All rights reserved.

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • Home
  • AI Research
  • Platforms
  • Ethics
  • Developer AI
  • Industry
  • Data Science
  • Emerging Tech
  • Quantum
  • BioTech
  • Decentralized

© 2026 Future News 24. All rights reserved.

Website security powered by MilesWeb