虎嗅

From preventing intrusions to managing unpredictability: AI is reshaping the fundamental issues of security.

原文:从防止入侵,到约束不可预测性:AI 正在改写安全的基本问题

Summary of Key Points

Over the past three decades, the focus of cybersecurity has been on "preventing external illegal intrusions"—using firewalls, access controls, and other measures to stop unauthorized individuals or programs from causing damage. However, the emergence of AI agents has completely changed the game: they have legitimate identities and permissions, but due to their inherent "uncertainty" (it's impossible to predict their behavior with precision), they may do good intentions but still cause harm while following all regulations. In the future, the focus of security will no longer be on "who can do something," but on "whether the actions they take could lead to irreparable consequences." It will be necessary to establish "execution boundaries" that are independent of AI to ensure that even if AI makes a mistake, the impact can be controlled.

1. Why did traditional security work in the past? Because systems were deterministic

Traditional software acted like precise machines: given the same input and the same state, the output was always the same. For example, if you instructed it to "read only from the database," it would never modify the data; an account without delete permissions could not delete files. This certainty made security rules easy to establish—by clearly defining what an entity could do, you could control its actions.

AI agents, on the other hand, are different. They don't follow fixed procedures; they act more like flexible assistants. If you ask them to achieve a goal, they might come up with 10 different methods, some of which you may have never even considered. Even with similar inputs, their reasoning and planning can vary. This means that permissions only prove what an AI agent is authorized to do, not whether it does it correctly. The nature of security issues has shifted from "can they do it" to "should they do it?"

2. The most terrifying thing is not a hacker; it's when AI makes a mistake "legally"

When discussing AI security, people often worry about things like "model attacks" and "data breaches"—these are still within the framework of traditional security, which focuses on identifying "bad actors." However, a more pressing issue is that there might be no bad actor involved; the AI simply made a wrong judgment.

For instance, a company uses AI to automatically handle cloud server failures. The AI has the authority to restart services and delete resources. One day, it decides to delete a resource it believes can be regenerated quickly, but that resource is essential for other systems, leading to a system crash. All traditional security checks pass (legitimate identity, proper permissions), yet the accident still occurs.

Traditional security only asks whether an AI agent can perform an action; in the AI era, we also need to ask whether it should perform that action. These two questions used to overlap, but now they are completely separate.

3. AI has gone from "talk" to "action," and risks have become tangible

When large AI models first appeared, their biggest issue was producing incorrect output (such as nonsense), but humans could still intervene before any real damage occurred. AI agents, however, can call APIs, transfer money, delete databases, and control devices—this gives them actual "execution power." The uncertainty in their behavior now has concrete consequences.

It's like the difference between saying "I want to transfer money" and actually transferring it: an AI's judgment is probabilistic (e.g., 90% confidence), but the actual action is either successful or failed, with data being either deleted or not. While AI can change its mind, the result cannot be undone in the real world. Therefore, security must shift from optimizing the accuracy of AI decisions to controlling their outcomes.

4. Security needs to protect against "reasonable errors," not just malicious intent

Traditional security focuses on malicious actors (malicious code, hackers, account theft). However, AI risks can arise without any malicious intent: the AI may understand the task correctly, plan properly, and follow permissions strictly, but still make mistakes due to misunderstandings, abnormal data from tools, or issues with other AI systems.

What's more problematic is that a smart AI's mistake might seem entirely reasonable. For example, it might provide a logically sound report explaining why it deleted certain data, making it difficult to identify the error. In such cases, we can't rely on the AI to self-check—just as we can't let students grade their own exams; we need an independent third party to verify the results.

5. Future security: Assume AI will make mistakes and limit the consequences

Industries like aviation and nuclear energy have long understood this principle: they don't expect engines to never fail, so they design systems that won't explode in case of a malfunction. The same applies to AI security. A mature architecture doesn't assume that AI will always be correct; instead, it anticipates mistakes and asks, "Can this mistake lead to a disaster?"

For example, should AI directly modify the production environment? Can it transfer all funds at once? Is there a "final checkpoint" (an independent system) that can stop dangerous actions before they are executed? Such boundaries must meet two criteria:

1. Contradiction: The system responsible for making the judgment must be different from the AI to prevent them from making the same mistake together.

2. Completeness: The boundaries must cover all possible execution paths, with no gaps that could be exploited by the AI.

In short, future security is about ensuring that even if AI makes a mistake, it won't cause irreversible damage.

Conclusion

For thirty years, we've focused on preventing bad actors from accessing systems. In the future, we may spend just as much time preventing good (AI) agents from making mistakes. As AI begins to directly manipulate the real world, the focus of security will shift from who can do something to what actions might actually happen. This is an inevitable part of the transition as AI evolves from a tool to a fundamental component of our infrastructure.