虎嗅

"Hidden Reasoning and Explainability in Large Models: What is AI Thinking? | Interview with Stanford Professor Aryaman Arora"

原文:大模型的隐藏推理与可解释性:AI在想什么?|对话斯坦福博士Aryaman Arora

Summary of the Key Points

This news article focuses on the mystery of the “black box” of AI: Through an experiment that manipulated the internal cognition of an AI model regarding animals (changing the answer from 8 to 6), it reveals the hidden internal reasoning processes that AI uses when answering questions. The article then discusses how researchers are trying to address critical questions regarding the “black box” (such as whether the thought processes represent genuine thinking and whether the model truly understands the world). It also mentions new developments in AI, such as the ability to adjust responses based on users and the emergence of “personalities” among AI models. Finally, it emphasizes the importance of explainability in building trust and making decisions regarding whether to entrust important tasks to AI.

Detailed Explanation

1. Hidden Reasoning Behind AI Answers

The experiment described in the article is quite illustrative: When asked how many legs animals that can weave webs have, the AI initially answered 8, assuming the animal was a spider, even though the question did not explicitly mention one. Researchers simply replaced the internal assumption of the model from a spider to an ant, and the answer changed to 6. This indicates that AI makes a key judgment before providing an answer—first identifying the type of animal and then determining the number of legs. This process is similar to the “inner assumptions” humans make while thinking, but AI does not reveal these steps, which is why it’s called “hidden reasoning.”

2. The Difficulty of “Opening the Black Box”

We know that AI has some form of internal reasoning, but we don’t yet understand the exact steps it follows to reach its conclusions. For example:

  • Are the thought processes (the sequence of steps AI uses to arrive at a answer) a true representation of its thinking, or are they merely imitations of human thought?
  • Does the model truly understand concepts like “spiders weaving webs” or “ants having six legs,” or does it merely memorize statistical correlations?

Researchers have attempted to crack these mysteries by observing the model’s internal workings and manipulating variables, but there are still no clear answers, and the black box remains largely closed.

3. Understanding AI’s Correct Answers Is Easier Than Its Mistakes

The article points out a counterintuitive phenomenon: It is often easier to explain why AI gets things right than why it gets them wrong. For instance, when AI correctly calculates 1+1=2, we can understand that it is using the addition rule. However, if it calculates 3, the issues might be due to incorrect parameters, confusion of knowledge points, or interference from other data, making it harder to identify the error. Similarly, when a human gets a question right, they can usually explain their reasoning; when they get it wrong, they may not even be aware of where they went wrong.

4. New Trends in AI

There are two notable developments in AI:

  • Adaptive Responses Based on Users: AI now adjusts its answers to the user’s context. For example, it may use simpler language with children or more technical terms with experts.
  • Diverse Personalities: Different AI models exhibit different personalities, ranging from lively to rigorous, or even sarcastic.

These changes make AI more “human-like,” but they also require a better understanding of its behavior—such as why it treats different people differently or whether it exhibits any biases.

5. The Importance of Explainability

The reason for studying explainability is twofold:

  • Building Trust: When using AI for tasks like medical diagnoses, we need to be confident that its conclusions are based on reliable medical data, not on random guesses.
  • Avoiding Risks: For example, if AI is used for loan decisions, we need to understand the reasons behind its decisions to ensure there is no bias against certain groups.
  • Setting Boundaries: We need to determine which critical decisions can be entrusted to AI. In areas like autonomous driving and financial risk management, it is essential to understand the logic behind its judgments before relying on it.

As AI becomes more capable, understanding its “thinking processes” will become increasingly important. After all, we cannot entrust our fate to a completely opaque “black box.”