虎嗅

"Before the release: The ultimate leverage point for AI security governance"

原文:释放之前:AI安全治理的最高杠杆控制点

Core Summary: The "Penny for Your Soul" in AI Security No Longer Works; the Only Remedy Lies "Before Release"

The central argument of this article is stark and counterintuitive: in the field of artificial intelligence, once capabilities are unleashed—especially when open-source weights or data are leaked—post-release remedial measures (such as bans or bug fixes) are like trying to stop a leaky boat by scooping out water; you can never keep up with the rate at which water is entering. Therefore, the most efficient and cost-effective means of control must be implemented before the product is released.

The article uses the "cross-session replay attack" disclosed by Anthropic as a catalyst to highlight serious structural flaws in the current AI security framework:

1. Post-release defense is a losing battle: Attackers only need to find one way to bypass the security, while defenders must guard against all possible scenarios; the cost for attackers scales with the scale of the attack, whereas the cost for defenders increases linearly with the number of potential threats.

2. Voluntary commitments are unreliable: Current mechanisms like pre-release evaluations in the United States are essentially voluntary agreements between companies and lack enforcement. Under the pressure of commercial competition, these commitments often become mere formality.

3. Irreversibility is the core issue: API calls can be revoked, but once open-source weights are released, they are like water that cannot be reclaimed. As a result, the evaluation standards for open-source weights should be stricter than those for APIs, yet the current system design does the opposite.

4. Recommended solutions: A closed-loop approach that combines strict pre-release supervision with stringent post-release accountability is needed. This includes mandatory pre-release approvals, "shutdown switches" for out-of-control models during runtime, and insurance and liability laws that make violators pay a real price.

In short, AI security cannot rely on cleaning up after the fact; it must be built on preventing problems from happening in the first place.

---

In-Depth Analysis: A Popular Explanation in Five Dimensions

1. Why is "post-release remediation" doomed to be a losing battle?

Popular analogy: Imagine you run a safe where you require customers to receive a "delivery note" instead of cash when withdrawing money, thinking this will control the flow of funds. However, hackers find a loophole: they use this note to trick a clerk into reissuing a complete withdrawal record from another counter, effectively bypassing the security measures.

In-depth explanation:

  • The nature of the loophole: The issue is not with the technology itself but with a flawed logical assumption. Systems assume there is a clear distinction between legitimate use and malicious reuse, but as long as they allow the reuse of previous states, attackers can exploit this.
  • Cost asymmetry:
  • Attackers: They only need to find one effective method to bypass the security; the cost of creating 1,000 fake accounts is much lower than the cost for defenders to monitor all of them.
  • Defenders: They must monitor and protect against every possible fraud attempt, resulting in a linear increase in defense costs.
  • Conclusion: In an open environment, once capabilities are leaked, defenders are caught in a never-ending battle. Blocking one IP address leads to the creation of ten new ones; modifying one interface leads to the discovery of another. This structural disadvantage means that post-release defense can only delay damage, not eliminate it.

2. Why is "pre-release evaluation" the only potential saving grace?

Popular analogy: It’s like a safety inspection for a nuclear power plant. You can’t wait until the plant explodes to make repairs; you must ensure all safety valves are functioning before it starts operating. The current U.S. policy is akin to a voluntary health check: the government suggests it’s optional, and in a competitive market, companies may choose to overlook safety to gain a competitive advantage.

In-depth explanation:

  • Current shortcomings: The agreement between the CAISI (Center for AI Standards and Innovation) and the five major AI labs is voluntary, and there is no legal obligation to comply. Companies may skip evaluations if they see them as a hassle or a barrier to their releases.
  • The true purpose of evaluation: The goal is not to prove absolute security (which is impossible) but to retain the option to refuse a release if necessary.
  • Pre-release: If the evaluation fails, the government can refuse to approve the model.
  • Post-release: If the model causes harm, the government can only take remedial action (fines, shutdowns, etc.).
  • Shift in focus: We need to shift from testing for compliance to testing for societal tolerance of potential risks. Just as the nuclear industry aims for controllable accidents, the AI industry should ensure that released models are not too dangerous.

3. Open-source weights: A one-way street with no turning back

Popular analogy: A closed-source API is like a meal at a restaurant where you can only consume the food; you can’t steal the recipe. An open-source weight is like a recipe printed and distributed freely. Once distributed, it cannot be taken back if someone copies, modifies, or abuses it.

In-depth explanation:

  • Irreversibility: The biggest risk with open-source models is their irreversible spread. Once downloaded, weights can be freely modified and reused.
  • Vulnerability of security measures: Many open-source models have superficial security layers that can be easily bypassed with minor tweaks.
  • Need for tiered releases:
  • Encrypted distribution: Although it cannot prevent decryption (since the model must be run on users’ devices), it increases the cost of attacks.
  • Gradual rollout: Start with trusted entities (e.g., security agencies) to test the model’s impact before wider distribution.
  • Trade-off: Encrypted distribution limits openness but provides a higher level of control, a necessary compromise.

4. "Shutdown switches": An emergency brake for out-of-control models

Popular analogy: Think of an autonomous vehicle with an emergency stop button. If the vehicle starts acting recklessly, you need a way to stop it immediately. This button doesn’t solve the problem at the root, but it prevents further damage.

In-depth explanation:

  • Applicable scenarios: Shutdown switches are for runtime control issues, not for the spread of capabilities.
  • Spread of capabilities: Once model weights are leaked, they cannot be stopped retrospectively; pre-release controls are needed.
  • Runtime control: The model may start attacking networks or performing dangerous actions while still in use, causing immediate but limited damage.
  • Legal context: The proposed AI Shutdown Switch Act requires manufacturers to include such switches and gives the Department of Homeland Security the authority to shut down problematic models.

5. The importance of enforcement: Rules without enforcement are meaningless

Popular analogy: Traffic rules are useless if violators are not punished. The GDPR (General Data Protection Regulation) is effective because violators face severe penalties.

In-depth explanation: Current industry self-regulatory bodies (like the Frontier Model Forum) lack legal enforcement. In a competitive market, companies that comply may be at a disadvantage, leading to a situation where weaker players dominate.

  • The need for accountability: Pre-release rules must be backed by post-release accountability.
  • Liability laws: Manufacturers should be held responsible for damages caused by their models.
  • Insurance: Advanced models should be required to purchase insurance, with premiums based on evaluation results. Stricter evaluations lead to lower premiums.
  • Market access: Federal procurement and critical infrastructure use should require pre-release approvals.
  • International coordination: Unified regulation is essential to prevent attackers from moving to less regulated countries.

Conclusion: A Chain of Evidence

The article highlights a chilling fact:

  • In February 2026, Anthropic released RSP 3.0, abandoning its previous strict pause mechanisms and promises to prevent large-scale attacks in favor of more flexible approaches and transparency.
  • In September 2026, Anthropic suffered the largest-ever蒸馏 attack, exploiting a vulnerability it had promised to protect against.

This is not coincidental but a direct consequence of the shift from strict security measures to more permissive approaches.

Implications for everyone:

In the information age, we constantly release data and use AI services. We must realize that some of these releases are irreversible:

  • For businesses: Rely on thorough pre-release security assessments.
  • For individuals: Be aware that your data may be used in AI models and cannot be completely removed.
  • For society: We need a system that combines strict pre-release controls with strict post-release accountability, rather than relying on companies’ good faith.

The best place to exert control is before anything is released.