When the “Safeest” AI Company Begins to “Run Naked”: A Game of Trust, Money, and Survival
Hello everyone, I’m your financial journalist. Today, we’re not talking about a dry technical report, but about a “crisis of trust” that’s happening at the top of Silicon Valley.
The protagonist of this story is Anthropic, a company widely regarded as the most rule-abiding and safety-conscious in the AI industry. The key figure is Jacob Coxon, a researcher who had only been with the company for four months. On September 8, 2026, he suddenly resigned and publicly stated that both Anthropic and OpenAI are risking human lives in their pursuit of superintelligence, abandoning all safety precautions.
What made this incident so explosive wasn’t just the resignation itself, but the truth Coxon revealed: “I left before the equity was allocated; I no longer needed to boost Anthropic’s valuation for my own benefit.”
Imagine you have a huge lottery ticket that’s about to be cashed in (Anthropic is expected to go public in October with a valuation of $2 trillion), but you decide to tear it up. Why? Because you believe the company is doing something incredibly wrong.
Today, we’ll break down the logic behind this with simple language to understand what’s really going on.
---
From “Brake Pads” to “Decorative Strips”: The Secret Revision of Safety Commitments
First, let’s clarify what exactly changed at Anthropic.
In 2023, Anthropic released the Responsible Expansion Policy (RSP), which contained a very strict commitment that could be compared to a car’s “emergency braking system”:
“If our safety measures are not in place, we must pause the training of more powerful models.”
This was a hard and absolute rule that distinguished Anthropic from OpenAI, which focused on speed.
However, in February 2026, Anthropic released RSP 3.0. This update removed that “emergency brake” and replaced it with a more complex “dashboard” consisting of roadmaps, risk reports, and external supervision.
In simpler terms:
Before, it was “If the brakes fail, I stop the car.”
Now, it’s “If there’s a problem with the brakes, I write a report explaining why I still drive, and I drive while monitoring the dashboard.”
Anthropic argued that since the boundaries of AI’s capabilities are too vague to define precisely when to stop, a rigid brake system is impractical. This makes sense, but the question is: If the boundaries are so unclear, why was this commitment such a highlight just before the IPO?
Even more concerning is that around the same time, Anthropic signed a $517 billion contract for computing power for the next decade. A company that locks in its future expenses for a decade is less likely to actively apply the brakes. The safety commitment shifted from “what we don’t do” (a hard rule) to “how we explain what we do” (a soft explanation).
Two Words That Disappeared: OpenAI Also “Removed Their Makeup”
Do you think only Anthropic was changing its policies? No, OpenAI was doing the same.
In its annual tax filing (Form 990) in November 2025, OpenAI quietly revised its mission statement:
- Previously: “To build safe, general-purpose artificial intelligence that benefits all of humanity, unrestricted by financial profit considerations.”
- Now: “To ensure that general-purpose artificial intelligence benefits all of humanity.”
Two words disappeared:
1. “Safe” (the commitment to safety was downplayed.)
2. “Unrestricted by financial profit” (the commitment to independence from financial interests was removed.)
This means OpenAI no longer emphasizes that it exists for human safety or that it is independent of capital interests. It now behaves more like a pure profit-maximizing tech giant.
Meanwhile, OpenAI disbanded three of its core safety teams within two years, including the renowned “Super Alignment Team.” Even a visionary like Ilya Sutskever left to found a company dedicated to safety.
Interpretation:
These two companies, which once emphasized safety the most, are both redefining the boundaries of safety at the same time. This isn’t just a moral decline for one company; it’s a collective shift in the entire industry under the pressure of capital and competition. When “safety” becomes a negotiable variable rather than an unbreakable principle, danger arises.
An Unsolvable Paradox: To Be Safer, You Must Be Faster?
Coxon made some stark statements in an interview that reveal the inherent dilemma of the AI industry:
1. “If you’re under competitive pressure, you have to take shortcuts or skip safety checks.”
2. “Excessive paranoia towards OpenAI and China is sometimes used as a justification for accelerating progress.”
This creates a terrifying safety paradox:
- Logic A: If Anthropic slows down for safety assessments, OpenAI or Chinese competitors may develop superintelligence first.
- Logic B: If competitors succeed first and their safety measures are weaker than Anthropic’s, global safety risks increase.
- Conclusion: To be “safer” (to prevent unsafe entities from developing superintelligence first), Anthropic must develop superintelligence faster.
- Consequence: But being faster means shorter training, assessment, and release times, reducing the time researchers have to understand model behavior and identify risks.
In plain language:
“To be safer, I must run faster. But the faster I run, the less safe I become.”
Coxon resigned not because he hated Anthropic, but because he realized that no company can maintain safety in this competitive environment. It’s a systemic dead-end. Changing companies is useless; the entire race is accelerating, and no one dares to apply the brakes, as they would be eliminated or overtaken by their competitors.
Three Days of Shutdown: Safety Depends on “Firefighting After the Fact,” Not “Prevention”
The article mentions a key event: the Fable 5 incident in June 2026.
Anthropic released the new Fable 5 model, and two days later, the Amazon CEO called the White House, pointing out a vulnerability that could allow the model to bypass security measures. The White House convened urgently, and the Commerce Department issued an export ban, resulting in the global shutdown of Fable 5.
From release to shutdown, it took only three days.
A harsh reality:
What stopped the model wasn’t Anthropic’s revised safety rules; it was the government’s intervention after the risk had already occurred.
This shows that Anthropic’s safety commitments failed to act as a preventive measure. They were more like a post-event public relations tool or a form of compliance. When the risk became significant, external forces (the government) intervened within three days. However, the problem had already arisen.
The revised RSP 3.0 was supposed to work before the risk occurred. Now, it’s ineffective. What we see is not prevention but post-incident handling.
Investors’ Nightmare: How Much Do Safety Commitments Worth?
Finally, let’s look at the capital market.
Anthropic’s valuation is $2 trillion, and its brand narrative is “We are the safest AI company.” Investors buy not only the model’s capabilities but also the promise of safe governance.
However, if safety commitments can be altered under competitive pressure, they should be discounted in investment decisions.
- A Morningstar report states that the Fable 5 incident turned safety commitments into financial risks.
- David Sacks (an American tech advisor) advocates for a pause in Anthropic’s IPO until the situation is clarified.
What does this mean for ordinary people?
1. Brand Premium Collapse: If “safety” is just marketing hype and not a real commitment, Anthropic’s valuation is based on nothing solid.
2. Compliance Crisis: If regulators find discrepancies between its claims and actions, it’s a crisis of both public relations and legal compliance.
3. Vote of Trust: Coxon’s resignation is a vote of distrust. He used his own money to tell the market: “I don’t trust this system’s ability to correct itself.”
In summary:
Coxon’s resignation may seem insignificant compared to a $2 trillion valuation, but his words are a critical warning.
They punctured the AI industry’s inflated bubble: “We think we’re building a safer future, but in reality, we’re creating more expensive, uncontrollable machines that rely on government intervention.”
When “safety” becomes a matter of explanation rather than a firm principle, we lose not just technical safeguards but also moral foundations.
For the public:
- Don’t blindly trust tech giants’ safety claims; instead, watch their actions and regulatory responses.
- Pay attention to AI industry regulations, as current safety relies heavily on government intervention rather than company self-discipline.
- Be wary of a culture that values speed above all else: In AI, speed often comes at the cost of safety, and the price may be paid by all of humanity.
There are no winners in this game. Anthropic has lost its visionary researcher, OpenAI has lost its moral high ground, and investors stand on the brink, watching a car without brakes race ahead.