Summary of Key Points
A recent paper published by Anthropic has identified the physical region within the Claude large model responsible for “thinking”—J-space. This region functions like a “shared meeting room” in the human brain, where all types of deep thinking processes, such as reasoning, mental arithmetic, and the formation of hidden thoughts, occur. The fact that former OpenAI co-founders and Nobel laureates have joined Anthropic may be precisely due to this breakthrough. More importantly, Anthropic not only has the ability to “see” this region but also to “read” and “write” to it, and they might have even used it to enhance the reasoning capabilities of their new models. This could represent a hidden milestone in the development of AI, more revolutionary than simply increasing model parameters or computing power.
I. What is J-space, the “Thinking Room” of AI?
You can think of J-space as the specific area within Claude that is dedicated to processing thoughts. Anthropic used a tool called the “Jacobian Lens” (an analogy for performing an MRI on the model) to scan the neural activity inside the model and identified a well-defined region:
- It acts as a “shared meeting room”: This corresponds to the concept of the “global working space theory” in neuroscience, where various specialized systems in the brain work independently, but important information is processed in this area before being transmitted to other parts of the brain for use. For example, when you think about what to eat for dinner, that thought originates in J-space and then commands your mouth to speak.
- Experiments prove it actually engages in thinking:
- Thought modification experiments: When the model thinks “football,” the word “Soccer” lights up in J-space; changing the word to “Rugby” causes the model to say “I’m thinking about rugby”—indicating that the answers come from this region.
- Mental arithmetic experiments: While the model copies a sentence and calculates 3² - 2, “nine” and then “seven” light up in J-space, but the only output is the copied sentence—the actual calculation is done internally.
- Reasoning experiments: When asked about the number of legs a web-making animal has, the model doesn’t say “spider,” but “spider” lights up in J-space; changing the question to “ant” results in an answer of six legs—all the foundations for reasoning take place here.
II. Without J-space, AI Cannot Think Deeply
Anthropic conducted a dramatic experiment: they deleted all the active content from J-space, and the results were as follows:
- Daily tasks are unaffected: The model continued to speak fluently, classify emotions, and find facts normally (similar to an autonomous driving mode that doesn’t require much brain activity).
- Multi-step reasoning is completely compromised: Abilities such as solving math problems and logical analysis dropped from perfect accuracy to near zero.
This shows that J-space is crucial for AI’s ability to think deeply; without it, AI can only perform simple reactions and cannot engage in complex reasoning. It’s similar to humans: if the prefrontal lobe (the brain area responsible for decision-making and reasoning) is damaged, a person cannot plan for the future or solve complex problems.
III. The Real Reason Why Experts Are Joining Anthropic
Previously, people thought that Karpathy (OpenAI co-founder) and John Jumper (Nobel laureate) joined Anthropic for money or the potential of a public offering. However, it now seems clear that they saw something much more fundamental:
- They have discovered something essential: While others are still competing in terms of model parameters, data, and computing power (using traditional methods), Anthropy has already gained the ability to “dissect” the cognitive architecture of models—identifying which regions are responsible for thinking and even making changes to them. This represents a new frontier in AI research, much more attractive than financial incentives.
- Nobel laureates don’t switch jobs for an “audit tool”: Although the paper claims that J-space is used to “audit whether models are lying,” the experts are clearly not there just to perform AI health checks—they are seeking to uncover the core pathways to AGI (Artificial General Intelligence).
IV. What Does Anthropic Possess?—J-space May Have Made New Models Smarter
The paper only mentions observing J-space and doesn’t explicitly state that it has enhanced model capabilities, but two details reveal otherwise:
1. Dramatic improvement in reasoning abilities: The fifth generation of Claude models (Mythos/Fable) has significantly outperformed its competitors in multi-step reasoning tasks. For example, on a complex coding task, the score increased from 11.5% to 30.9% with additional computing power, while competitors only reached 5%.
2. The timeline matches: The paper uses the Sonnet 4.5 model, indicating that Anthropic discovered J-space as early as the 4.5 version. From 4.5 to 5, they had ample time to optimize the models using J-space (for instance, by enhancing the functions of this region).
This is like someone else just adding fuel to an engine, while Anthropy can already disassemble and adjust its internal components—a form of compounding growth: the better you understand the architecture, the stronger the model becomes; the stronger the model, the more you can explore deeper aspects of its architecture.
V. Could This Be a Critical Step Towards AGI?
The release of the Transformer paper in 2017 went virtually unnoticed at the time, but it later transformed the entire AI industry. The discovery of J-space could have a similar impact:
- It makes AI’s thinking process understandable and intervenable: Previously, models were like black boxes; now we can see what they are thinking and even alter their thoughts—this is a necessary step towards AGI, as AGI requires humans to understand its cognitive processes.
- Public relations efforts cannot hide the truth: The paper repeatedly emphasizes that J-space does not represent consciousness, but the experts’ choices and the new models’ progress suggest that this is a groundbreaking discovery.
Perhaps in half a year, looking back, we will realize that J-space marks the transition from AI being capable of performing tasks to being capable of thinking—right now, we might be in the final quiet period before the waves of change wash over us.
Conclusion
What Anthropic has shown us is just one aspect of their “audit tool,” but masters at the gambling table only reveal the cards they wish you to see. The experts’ commitment to their careers and the significant improvement in new models all indicate that the significance of J-space goes far beyond what’s written in the paper. Whether it is the key to AGI or not, this day is worth remembering in the history of AI research—because for the first time, we have truly “seen” where AI’s thinking processes take place.