虎嗅

Within three years, AI control will move towards the forefront of safety measures.

原文:三年内,AI 控制会走向安全的一线

Core Summary of the Article

The main argument of this article is that AI is transitioning from a phase of "capacity expansion" to one of "control and constraints." Currently, people are focused on what AI can do (such as writing code or processing documents). However, with the emergence of AI agents—AI systems that can directly execute tasks—the focus has shifted from what AI can say to what it can actually do. The security risks have evolved from mere content errors to potential harm resulting from actual actions. The key to future AI development is not to make AI more powerful, but to establish a systematic set of control mechanisms that clearly define what AI can and cannot do, enabling it to safely integrate into critical areas such as corporate core systems and finance.

Detailed Analysis

1. AI's Growing Challenges: From “Showing Off Skills” to “Following Rules”

Society's expectations for AI currently revolve around its ability to perform tasks—writing code, providing customer service, processing data, etc. This is similar to how people were concerned about the power of steam engines or electricity when they first emerged, wondering how many machines they could drive or how many cities they could illuminate. However, all technologies eventually move from focusing on what they can do to determining what they should do. For example, when cars first appeared, people only cared about their speed; later, traffic rules and traffic lights were necessary. AI has reached this turning point: having the ability alone is no longer enough; we need to establish clear guidelines for its use.

2. AI Agents Bring Risks to Life: From “Giving Suggestions” to “Taking Action”

Previous versions of AI acted more like advisors, answering questions or generating content, and errors could be easily corrected. But AI agents are different; they can directly interact with real systems—calling payment APIs to transfer funds, deploying code in production environments, modifying corporate databases, etc. This is akin to asking a friend if they want to transfer money and then receiving the transaction themselves; the risk has shifted from incorrect information to actual financial losses. The question is no longer whether AI can do something, but whether it is authorized to do so.

3. Changes in AI Security: From “Correctness of Content” to “Appropriateness of Behavior”

Conversations about AI security often focus on whether AI generates illusions or incorrect information. In the future, the focus will be on behavior control. For instance, if AI can generate code, should it be allowed to deploy it directly? If it can analyze payment transactions, should it be authorized to initiate them? If it can access data, should it be allowed to delete logs? Any mistake in these actions could result in significant financial losses or system failures. The same AI can pose vastly different risks depending on the task; it is essential to clearly define what it can and cannot do.

4. Traditional Permissions Are No Longer Enough: Customized “Traffic Rules” for AI

In traditional software systems, permissions were assigned to humans—accounts belonged to people, and approvals were made by humans. However, AI agents execute tasks autonomously. In this case, relying solely on account permissions is insufficient; we need to ask whether a particular action is appropriate for the given context and whether the AI understands the associated risks. If the AI is manipulated, can the system prevent it from making harmful decisions? For example, if AI is about to transfer a large sum of money, we need to consider not only whether it has the permission but also who the recipient is, why the transfer is happening, and whether the amount exceeds any risk thresholds; human approval may be required.

5. A Control System Is the Key to Accessing Critical Areas

For AI to enter critical fields such as finance, corporate core systems, and industrial control, a robust control system is essential. Just as a car without brakes cannot be on the road, an aircraft without air traffic control cannot fly, and a financial system without risk management cannot conduct transactions. The competition in the future will not solely depend on which AI models are smarter; it will also depend on who can ensure that AI operates safely. Defining the boundaries of AI will determine who can integrate it into core business processes confidently. This is not about limiting innovation but about creating the conditions for AI to truly make a difference.

In One Sentence

AI is becoming increasingly capable, but being capable does not mean it can do anything it wants. The focus should now be on establishing rules and boundaries to ensure that AI operates within a safe framework, enabling it to become a powerful tool for transforming the world.