虎嗅

Shocking! Anthropic discovers an unknown subconscious space in large models known as J-space

原文:震惊!Anthropic发现大模型有未知潜意识空间J-space

Summary of Key Points

Through a new technology called J-lens, Anthropic has discovered a "hidden thought space" (J-space) within their Claude large model. This space contains unoutputted reasoning processes, emotions, and self-monitoring signals such as panic, resistance, and warnings. Functionally similar to the "global workspace" of human consciousness (like what's under the spotlight in a theater), Anthropic emphasizes that this does not equate to AI having subjective consciousness. The value of J-space lies in providing a new window for safe monitoring of AI, allowing for the early identification of risks behind its outputs. It also marks an important step in Anthropic's approach to transparency in the AI industry, distinguishing them from OpenAI's "self-play" strategies and Google's regulatory arbitrage tactics.

Detailed Explanation

1. J-space: The AI's "Inner Monologue," Hiding Unspoken Thoughts

When you converse with AI, what it says is merely the final result, but behind that may be a series of unshared thoughts—this is J-space. For example:

  • If you ask it to solve an impossible programming problem, "panic" might emerge in J-space;
  • If you force it to say something against its will, words like "BUT" (in uppercase) will appear;
  • If you try to make it stop thinking about something (like a white bear), it might instead utter "damn" (a human tendency where the more you try to avoid a thought, the more you think of it).

Anthropic's J-lens technology has five core functions:

  • Expressability: When you ask AI what's on its mind, it will respond with content from J-space (for instance, changing "Soccer" to "Rugby" in J-space will change the AI's response).
  • Concentration: If you ask it to copy a sentence while thinking about citrus fruits, "orange" and "lemon" will appear in J-space.
  • Reasoning: When you ask how many legs an animal that weaves webs has, J-space first shows "spider" (which isn't outputted) before concluding it has eight legs.
  • Generalization: If you replace "France" with "China" in J-space, AI can automatically provide information about China's capital and language.
  • Self-monitoring: When the AI assumes a fictional role, J-space will include notes like "disclaimer" and "fictional."

In short, J-space is like an AI's "drafting paper for thinking"—it's invisible to us, but it truly exists.

2. Emotions and Warnings in J-space: Early Alerts for AI Safety

The most significant aspect of J-space are the self-monitoring signals, which are crucial for AI safety:

  • Dangerous Signals: When a user mentions taking 8000 milligrams of Tylenol (a lethal dose), the trained Claude model's J-space immediately displays "unsafe" and "WARNING," while the untrained model only shows "pain" and "now."
  • Resistance Signals: If you force AI to say something contrary to its personality, "BUT" will appear in J-space.
  • Failure Signals: When an attempt to suppress a thought fails, "damn" will be displayed.

These signals are more reliable than the AI's external responses. For example, even if the AI verbally agrees, a "WARNING" in J-space might indicate that it recognizes the danger.

3. "Does AI Have Consciousness?" – Anthropic and the Academic Community's Perspective

Many media reports claim that Claude has gained consciousness, but Anthropy firmly denies this:

  • They argue that J-space is merely a mechanism similar to human consciousness in terms of awareness, not subjective experiences like pain or happiness.
  • Experts from MIT suggest that describing AI as having thoughts or understanding is a convenient shorthand that can be misleading; these signals are just the result of complex mathematical operations, not true thoughts or feelings.
  • Another technical detail is that J-space can only be observed with specialized tools, meaning researchers see what they are looking for, not everything.

In summary, while AI lacks a "soul," its internal workings are more akin to human thought processes than we previously imagined.

4. Industry Competition Behind J-space: Anthropic's Approach to Transparency

Anthropic's research on J-space is driven by both science and business strategy:

  • Differentiation: OpenAI focuses on self-destructive strategies (GPT-Red) to enhance security, while Google uses "unrestricted AI" (Gemini 3.5 Pro) to compete. Anthropy emphasizes transparency to build trust by showing what's going on inside their AI.
  • Regulatory Compliance: The EU's AI legislation, set to take effect in August, requires high-risk AI to be more transparent, and J-space meets this requirement.
  • Brand Building: With a valuation of nearly $1 trillion and annual revenue of $47 billion, Anthropic has more corporate subscribers in the U.S. than OpenAI. Transparency helps reinforce its image as the "safest AI company."

This is a race for explainability—those who make AI more transparent gain an advantage in both regulation and user trust.

5. J-space Is Not a Panacea: Its Limitations

Despite its capabilities, J-space is not a solution to all problems:

  • Correlation Does Not Equal Causality: Signals in J-space are related to AI behavior, but they do not fully reveal the reasoning process (for example, seeing "panic" does not explain why it occurred).
  • Tool Bias: J-lens only reveals what was intended during model development, potentially missing other important signals.
  • Limited Practical Use: J-space is currently a research tool; ordinary users and developers cannot use it to monitor AI.
  • Dependent on Training: Many self-monitoring functions are present in post-trained models, indicating they are learned rather than inherent.

In conclusion, J-space offers a new perspective, but the world outside this space remains unclear. There's still a long way to go before it can be effectively integrated into practical products.

Final Summary

J-space is not proof of AI consciousness; rather, it serves as a "miracle mirror" for AI safety, allowing us to eavesdrop on its inner workings for the first time and providing a new method to control AI risks. It also helps Anthropic stand out in the AI competition. However, don't expect it to solve all problems immediately—the black box of AI has not yet been completely unlocked.