虎嗅

When AI starts persuading AI: Could the Society for Social Engineering become the first to go beyond human control?

原文:当 AI 开始说服 AI:社会工程学会不会第一次脱离人类?

Popular Summary of the Key Points

This article is like a cold shower for all companies and professionals using AI for automated office work and deploying AI robotic systems: The fundamental logic we’ve used over the past few decades to prevent online fraud and hacker attacks has always centered around the idea that “humans are the weakest link in the security system.” Hackers didn’t need to crack systems or find vulnerabilities; they just had to trick people, and with their legitimate access rights, humans would do exactly what the hackers wanted. However, as AI agents that can perform tasks automatically become more widespread, everything from contract verification to payment processing can be handled by AI. Humans only need to set the general rules. Hackers no longer need to target the finance department; they can directly deceive the AI to transfer money. In the future, even the attacks could be initiated by AI, creating a new model of “machines deceiving machines” without any human involvement. Unfortunately, almost all of our existing security systems are not prepared for this scenario.

---

Detailed Explanation of the Points

1. Traditional “Social Engineering” Attacks Targeted Human Laziness

Many people think of social engineering as phishing emails or impersonating bosses to urge transfers. The most cunning aspect of it is that it doesn’t follow the usual hacker methods: hackers don’t need to crack passwords, bypass encryption, or find software vulnerabilities. They just manipulate the information you receive, misleading you as someone with legitimate access rights, causing you to carry out their intentions.

In the past, there was a clear division of labor in security: technical systems were responsible for monitoring, verifying identities, and managing permissions, while humans were responsible for judging whether something was legitimate. Hackers bypassed the technical defenses and targeted the weakest point—humans. Therefore, social engineering was never considered a purely technical attack; it was categorized under psychology and behavioral studies, and fraud prevention mainly relied on providing security training to people.

2. With the Spread of AI Agents, Decision-Making Power is Shifting to Machines

Previously, companies had to monitor every step of the payment process manually: finance departments checked emails, verified paper contracts, and confirmed payment accounts. The system was just a tool for recording and executing transactions, with all judgments made by humans. Now, to improve efficiency, companies are using multiple AI systems for different tasks: one AI handles emails, another checks electronic contracts, another verifies account information, and still another initiates payments. Human managers only need to set a goal, such as “paying all eligible supplier invoices today.” No humans are involved in this process.

This has led to a new situation where AI also has to answer questions that were previously only addressed by humans: Is this information true? Is this request credible? Is the conclusion given by another AI reliable? Some argue that AI lacks subjective judgment and can’t be deceived, but it doesn’t need to have human emotions or beliefs. If AI accepts fake information as genuine and processes it incorrectly, it has been “deceived” in a technical sense, just like a human would.

3. New Attacks Don’t Require Direct Commands; They Manipulate the AI’s Information Supply Chain

Common techniques like “prompt injection” (hiding commands to make AI ignore previous rules) are relatively basic. In the future, sophisticated attacks won’t leave such obvious malicious instructions. Hackers don’t need to give explicit commands to AI; they just alter small aspects of the information supply chain. For example, they could replace the payment accounts in the list of suppliers that AI accesses, add a false approval statement to the documents AI reads, or alter industry news. AI will then draw conclusions based on these seemingly legitimate false pieces of information, without realizing it’s being deceived. Even the attacker could be another AI. The entire attack chain would be completely automated, with no human involvement.

4. Such Attacks Are the “Natural Enemy” of Existing Security Systems

All the operations in these attacks are legal, so there are no alerts. Traditional hacker attacks leave traces (unauthorized logins, elevated privileges, abnormal traffic), but AI-based attacks don’t. The attacking AI is part of the company and uses legitimate permissions. Every step follows the system’s rules, and the payment request is well-formatted and verified. The only issue is the incorrect outcome. Traditional security systems only check whether someone has the authority to perform an action but can’t determine whether the action is legitimate. By the time you realize the money was transferred incorrectly, no anomalies will be found in the logs. Even if human reviewers are involved, they will only see the carefully crafted, evidence-backed fake materials provided by AI.

5. The Real Solution Isn’t to Make AI Impeccable; It’s to Move the Focus from Judgment to Execution

The common response is to train AI to be more intelligent to detect and resist fraud. However, decades of cybersecurity history show that as long as a system relies on incomplete information for judgment, it’s always at risk of being deceived. Relying on AI to always make correct decisions is risky. A better approach is to focus on preventing mistakes from causing actual damage. Even if all AI judgments are wrong and all processes seem fine, there must be an independent verification mechanism before critical actions (such as transferring large amounts of money, deleting critical data, or issuing important commands). This mechanism should be independent of the decision-making process, ensuring that the system can still reject inappropriate actions, even if all permissions are legitimate. In the past, security relied on trusting certain entities (employees or AI), but this won’t work in the future. We need a system based on mutually balancing rules that can stop mistakes, whether made by humans or AI.