Summary of Key Points
This article compares the governance logic of AI agents with the evolution of human social systems. Our current expectations for AI agents are still rooted in the ancient notion of seeking a “perfect entity”—one that is intelligent, aligned with human intentions, and free from error. However, the essence of modern systems is to acknowledge that entities are inherently imperfect and to design rules to mitigate the potential damage caused by their mistakes. The article argues that the true security of AI agents lies not in training them to be “always correct” but in establishing a framework that prevents their errors from leading to widespread real-world disasters, just as human societies rely on contracts, laws, and approval processes to ensure safe collaboration among ordinary people.
Detailed Analysis
1. Our Expectations for AI Agents Remain Traditional
In ancient times, people hoped to find wise rulers to govern their nations; today, our expectations for AI agents are similar: we want them to be more intelligent, more understanding, less prone to deception, and better able to resist manipulation. Essentially, we are looking for a “perfect entity” that will solve all our problems.
However, this approach has a significant limitation: it is not scalable. A small tribal community could maintain order with a virtuous leader, but as societies grow, it is impossible to ensure that every official is honest or every ruler is wise. The same applies to AI agents. When AI agents evolve from being capable of generating text to managing finances, modifying databases, and controlling devices, relying on their “perfection” alone is insufficient, as they will inevitably make mistakes (e.g., misunderstanding instructions, having incomplete information, or encountering unforeseen situations).
2. The Wisdom of Modern Systems: Focusing on Preventing the Spread of Errors
Modern societies function not because people have become perfect but because they have accepted the reality that humans make mistakes and have designed systems to mitigate these errors:
- Financial Approval in Companies: This is not a lack of trust in employees but a mechanism to prevent individual errors such as misappropriation of funds.
- Cross-Check in Aviation: It is not a lack of trust in pilots but a way to prevent operational mistakes.
- Dual Review in Banks: This is not a lack of trust in tellers but to avoid losses caused by a single error.
The goal of these systems is not to eliminate errors but to prevent them from destroying the entire system. For example, transferring a large amount of money requires approval from a leader, who can stop the transaction if there is a mistake; a pilot’s mistake can be caught by a co-pilot.
This shift represents a shift from relying on “good people” to relying on “good systems.”
3. The Special Risk of AI Agents: The Potential for Rapid, Viral Errors
Human errors are constrained by physical limitations (we need to sleep, walk, and can only do one thing at a time), but AI agents have no such limitations. They can quickly use multiple tools, modify numerous files, and spread errors to thousands of people. For instance, if an AI misunderstands an instruction to “remove unnecessary data,” it could delete an entire database in an instant, whereas a human might take hours to realize the mistake.
The real danger with AI is not that they will make mistakes (which humans also do), but that their mistakes can be executed on a large scale. If we continue to focus on making AI “more perfect,” we are essentially giving more power to an entity that can never be 100% correct, just as ancient societies relied on an idealized ruler.
4. The Key to Governing AI: Controlling the Power to Execute Errors, Not Eliminating Them
The article makes an important distinction: an AI’s “wrong thought” is different from its “wrong action.” For example, if an AI mistakenly decides to transfer a large amount of money but lacks the authority to do so, that is just a thought, not a disaster. If it mistakenly decides to delete server data but requires human confirmation before execution, no damage will be caused.
Therefore, the core of AI security is to prevent wrong ideas from becoming real actions. This can be achieved by setting limits (e.g., requiring human confirmation for transfers over a certain amount), implementing dual reviews for high-risk operations, and restricting direct access to sensitive data.
5. From Ethical Statements to Practical Rules: The Need for Concrete Regulations
In the past, discussions about AI ethics focused on abstract principles such as fairness and privacy protection. But now that AI can take action, these principles must be translated into concrete rules:
- “Do Not Harm Humans” → AI should not operate dangerous equipment (e.g., factory robots) without human supervision.
- “Protect Privacy” → AI should not send user data to external servers.
- “High-Risk Operations Require Supervision” → AI code must be reviewed by both automated tools and human engineers before deployment.
Ethics tell us what is right, but regulations tell us what the consequences of mistakes are. Only by turning abstract principles into enforceable rules can AI safely enter the real world.
Conclusion
The governance of AI agents must follow the path taken by human societies: from seeking a perfect entity to building reliable systems. We do not need AI to be error-free; we just need to ensure that their mistakes are within controllable limits. Just as modern commerce exists because contracts and laws handle breaches of agreement, the future of AI agents will be built on systems that can prevent errors, not on the idea of a perfect AI. This is what truly represents the “modernization” of AI agents.