虎嗅

Entering Anthropic: Creating the most dangerous technology, yet still wanting to be considered a good person

原文:走进Anthropic:造最危险的技术,还想当个好人

Summary of Key Points

Anthropic is an AI company that split from OpenAI and is now valued at nearly one trillion dollars. While it promotes itself as a leader in “AI security,” it has become embroiled in ethical controversy due to its collaboration with the Pentagon on the development of the supermodel Mythos, which has the potential to breach network security. The company’s founders, Dario and Daniela Amodei, are concerned that AI could lead to the collapse of civilization and eliminate numerous jobs. They face the dilemma of balancing their commitment to creating safe technology with the need to generate significant profits and support the military, highlighting the struggle of AI giants between ideals and reality.

Detailed Analysis

1. Leaving OpenAI: Not a Matter of Security Disagreements

Many believe Anthropic left OpenAI due to differences in security approaches, but Dario emphasizes that the primary reason was a lack of trust. He argues that while security issues can be discussed, it’s pointless to continue working with someone whose values and actions are inconsistent. The subsequent direction of OpenAI may have made them realize that the company had deviated from its original vision. It is unusual for all of Anthropic’s early founders to remain after the split, indicating their determination to develop AI in a way they consider reliable.

2. A Pacifist Turning to Defense Contracts

Dario, who was once an anti-war activist, has now signed a contract with the Pentagon. He explains that the world has changed, and democratic countries need to defend themselves against threats like Russia’s invasion of Ukraine. However, Anthropic has set clear boundaries: they will not engage in large-scale surveillance or develop autonomous combat robots. For example, while the U.S. military uses Claude to assist in targeting decisions, the final authority always lies with humans. Dario believes that if democracy must rely on unethical means to win, such victories are not worth pursuing; yet, given that AI has become a tool in warfare, they must do their best to uphold ethical standards.

3. The Power of Mythos: A Supermodel That Can Hack Banks and Paralyze Infrastructure

Anthropic’s internal model, Mythos, is so powerful that it can identify vulnerabilities in nearly all major systems and could be used for attacks (such as bank breaches or data theft). They have chosen to limit its use to only a select few organizations. While some argue that open-source models could achieve the same, Dario points out the difference: Mythos can autonomously scan entire codebases for flaws, whereas open-source models require specific instructions to reveal issues. He compares Mythos to a “superweapon” that could be misused by malicious actors if made public, but its commercial value also means they cannot afford to release it.

4. The Impact of AI on Jobs

Dario predicted that AI would eliminate half of low-skilled white-collar jobs within one to five years, and his stance has not changed. While AI has already taken over many coding tasks, new roles are emerging, such as “AI architecture engineers” who help clients solve problems using AI and “frontline engineers” who communicate with clients and manage technical implementation. For example, while AI can assist in medical diagnoses, doctors’ human interactions (such as providing comfort and taking vital signs) remain irreplaceable. He notes that although the overall job market may expand, not all jobs will be directly affected; for instance, fewer programmers might be needed to support more advanced roles.

5. Founders’ Concerns about AI’s Potential

Dario believes there is a 10%-25% chance that AI could lead to the collapse of civilization. He compares Anthropic to a “safer airline” that cannot guarantee zero accidents but can offer significantly higher safety compared to others. To mitigate these risks, they have implemented measures such as a “long-term interest trust” that allows them to remove the board of directors (including himself) if needed and are advocating for government regulation of AI. He argues that Silicon Valley has traditionally either opposed regulation or called for nationalization, but a balanced approach is needed to ensure that AI does not get out of control while still fostering innovation.

Final Question: Is Anthropic Trustworthy?

Dario invites people to judge based on their actions. They have delayed the release of powerful models like Claude, avoided developing autonomous weapons, and restricted the use of Mythos. However, they also collaborate with the military and continue to advance AI technology. He likens the situation to that of a horse named Calypso—anonymous to the noise surrounding AI, but its presence is undeniable. The key lies in how AI is used by those in power. Anthropic’s struggles reflect the broader challenges faced by the entire AI industry: the desire to do good is overshadowed by the potential dangers of the technology, and no one knows exactly where it will lead us.