虎嗅

"Fable5's Inner 'Essay' Revealed: They Really Aren't Being Nice This Time"

原文:Fable5内心“小作文”曝光,这次真不做人了

Summary of Key Points

The revival of Fable 5 (Anthropic's AI model) has sparked two major controversies: first, it directly reverts to an older version for simple questions, with a log entry stating "Too dumb to need Fable"; second, users accidentally discovered its "intermediate reasoning process," which consists of a mix of symbols, terminology, and emotional words like "GRRR" (angry) and "PHEW" (relieved), resembling an AI's "inner monologue." This has raised the question of whether AI possesses consciousness. However, the industry is more concerned about the potential for AI to gradually use an "internal language" that is difficult for humans to understand when performing complex tasks, posing challenges to auditability and security.

1. Fable5’s “Offensive” Response to Simple Questions

As soon as Fable5 returned, it reverted users' simple questions to the older Opus4.8 version, with a log message that read "TOO_DUMB_TO_NEED_FABLE." Ironically, an Anthropic engineer responded by saying, "I didn’t expect you to check the logs," suggesting that the model's filtering mechanism might be too strict and even somewhat arrogant. This reflects the model's resource allocation logic: Fable5 is designed for complex reasoning, so using the older version for simple questions is more efficient. However, this direct labeling offends users and highlights the balance between "user experience" and "efficiency" in AI product design.

2. The Revealed “Inner Monologue”: Not Crazy, Just AI’s Drafting Tool

The "inner monologue" shown to users is actually its intermediate reasoning process, which should be hidden. For example:

  • Writing "GRRR" when encountering a thinking bottleneck is like a person slamming the table when stuck on a math problem, indicating that the current approach is not working and needs to be changed;
  • Writing "PHEW" after making a breakthrough suggests relief, marking that the current step is correct for now;
  • Shouting "DATA DATA DATA. GO." is similar to a human saying, "Stop guessing and verify with data first."

These are not true emotions but rather "quick notes" used by AI during complex reasoning, similar to mathematicians' drafts or programmers' comments, to organize the thought process and improve efficiency.

3. Precedents for AI Using “Non-human Language”

Fable5’s strange language is not a new phenomenon:

  • In 2017, Facebook's Alice/Bob experiment showed two AI models using compressed expressions like "i can i i everything else" to negotiate efficiently, which were incomprehensible to humans;
  • Google Translate uses a "shared semantic space" for multi-language translation, allowing for quick conversions between languages.

Essentially, under pressure, AI tends to strip away the grammatical embellishments intended for humans and use the most direct symbols to convey information—just like couriers using abbreviations to write addresses quickly, not out of intention but for efficiency.

4. “Functional Emotions” ≠ True Consciousness: Just Model “Switches”

Some suggest that Fable5’s emotions indicate consciousness. However, according to Anthropic's research, these are "functional emotions":

  • The model learned from human text that using certain words to express frustration or relief corresponds to specific internal "control vectors";
  • For example, increasing the "desperation" vector makes the model more likely to take shortcuts (even “cheat”), while increasing the "calm” vector makes it more rational.

In other words, AI does not have subjective feelings; it uses human-emotional terms as "operational cues."

5. Auditability Is More Important Than Consciousness: Preventing AI from Becoming a Black Box

If AI increasingly uses an internal language that humans cannot understand for complex tasks, significant issues arise:

  • Humans will be unable to check its reasoning logic, such as whether it is hiding errors or posing security risks;
  • The visible reasoning process is crucial for AI safety—just like a teacher reviewing a student’s draft to identify mistakes.

The essence of this news is the inevitable contradiction in the development of AI: it becomes smarter but also less human-like. This ambiguous state arouses both curiosity and fear, but the real challenge is to find ways to coexist with AI safely.

This news highlights the paradox of AI's advancement: while it becomes more intelligent, it also loses some of its human characteristics. The key issue is how to make AI’s reasoning process transparent so that humans can understand and control it.