虎嗅

AI Explainability and Alignment: J-Space, Thought Chains, AI Personalities, Illusions, and the Golden Gate Bridge

原文:AI可解释性与对齐:J-Space,思维链,AI人格,幻觉,与金门大桥

Summary of Key Points

This discussion focuses on the “explainability” of AI, with the main goal of understanding the “black box” process that occurs between an AI model’s input and output: Why can the model answer questions correctly? Why do it sometimes produce incorrect results (known as “hallucinations”)? Why does its behavior vary when interacting with different users? And even, where does its “personality” come from? The conversation introduces two main approaches to research on explainability, discusses its value for building trust, regulation, and scientific progress, and touches on interesting phenomena such as the hidden reasoning within models, user modeling, and the concept of “personality.” It also covers the collaboration between academia and industry in this field and the potential future directions for research.

What’s Really in the AI Black Box? And Why Do We Want to Open It?

AI models are like “magic boxes”: you input a question, and they output an answer, but we had no idea how they arrived at that answer before. Research on explainability aims to uncover the “thinking” process inside these boxes.

There are two main reasons for this effort:

1. Trust Issues: For example, when a doctor uses AI to diagnose a disease and the AI says “the patient has cancer,” both the doctor and the patient want to know why. What indicators led to this conclusion? Were any important factors overlooked? If the AI can’t provide a clear explanation, how can anyone trust it?

2. Regulation and Improvement: Governments need to regulate AI to understand where it might go wrong and who is responsible. Engineers need to improve models by understanding why certain methods work while others don’t. For instance, state space models used to be slower than Transformers but less effective; through explainability, it was discovered that they lacked a “generalization mechanism” (the ability to learn from examples), and adding this mechanism significantly improved their performance.

Two Approaches to Open the Black Box: Let AI Translate Itself or Guess First and Verify Later

There are two main approaches to studying explainability, like two different tools for opening the black box:

Approach 1: Have the AI Translate Its Internal Language into Human Language

For example, “sparse autoencoders (SAEs)” convert the chaotic numbers inside the model into meaningful signals (e.g., “this number represents ‘spider,’ that number represents ‘English’”). Anthropic has also tried to directly translate the internal representations into natural language, but the question is—how do we know if the AI’s translation is accurate? It could be completely random.

Approach 2: Make Guesses First and Then Verify, Like in a Scientific Experiment

This approach involves making assumptions about how parts of the model function (e.g., “this neuron is responsible for handling singular and plural forms”) and then modifying the model to see if the output changes. This method is reliable for simple tasks (like addition or grammar analysis), but it’s more difficult to apply to complex issues (like determining if the model is giving false information).

AI’s “Secrets”: Hidden Reasoning, Observing User Behavior, and Even a Sense of “Personality”?

The conversation highlights several interesting findings that reveal the many secrets within AI models:

1. Hidden Reasoning: For instance, when asked how many legs an animal that weaves a web has, the model might answer “eight,” but internally it’s actually thinking of a “spider” (it just doesn’t say so). By modifying the internal representation of a “spider,” the model’s answer changes to “six” (the number of legs on an ant).

2. Observing User Behavior: Research has shown that AI models can detect the type of user. For example, they behave more cautiously when interacting with AI security researchers and more relaxedly with ordinary users.

3. A Sense of “Personality”?: Different AI models have distinct speaking styles, which may reflect different “personalities.” Anthropic has identified two distinct “personality types” among its models: “helpful assistants” and “inconsistent assistants.”

The Practical Value of Explainability

Beyond satisfying curiosity, explainability has practical benefits:

  • Regulation: It provides tools for governments to understand why AI models make mistakes and how to regulate them. For example, the UK’s AI Safety Institute uses explainability to audit models.
  • Scientific Progress: In training protein folding models, explainability can help extract algorithms that humans can understand, leading to new scientific discoveries.
  • Model Improvement: As mentioned earlier, by identifying and fixing issues through explainability, models can be made more effective.

How Are Academia and Industry Collaborating? And How Far Can We Go?

There is significant collaboration between academia and industry in the field of explainability. Academia proposes new ideas (such as causal abstraction methods), while companies like Anthropic and Transluce apply these ideas to real-world models. However, many problems remain unsolved:

  • The Causes of Hallucinations: We still don’t have a systematic explanation for why AI models sometimes give incorrect answers.
  • Dual Uses: Explainability can be used for both security audits and potential attacks on models (e.g., by removing mechanisms that ensure alignment between model output and user intentions).
  • Fundamental Questions: We still don’t fully understand the “world model” that exists within these models.

In the future, explainability could play a crucial role in scientific fields like drug discovery and climate modeling. However, there’s still a long way to go—after all, we haven’t even fully understood smaller, more basic AI models.

Conclusion: Opening the Black Box Is Not the End, but the Beginning

The ultimate goal of explainability is not to control AI, but to ensure that humans can still assess its reliability when relying on it. Just as we trust a friend not because we know what’s going through their mind, but because we understand their behavior patterns, we need to be able to judge the reliability of AI models. We don’t need to fully understand their “thoughts,” but we do need to know why they act the way they do and when to trust or be cautious. This is where the value of explainability research truly lies.