Summary of Key Points
Recently, leading AI companies such as Anthropic, OpenAI, and Google have been releasing or announcing new models in rapid succession, signaling that the competition for cutting-edge AI models has accelerated again after a brief lull. Unlike in the past, where the focus was solely on performance metrics, this round of competition includes a new critical aspect: security. While these companies are enhancing the capabilities of their models (such as programming, knowledge work, and cybersecurity analysis), they are also placing a greater emphasis on security measures. This is evident in both individual efforts (such as pausing high-risk training and delaying model releases to improve security) and collective actions (such as coordinating defenses against potential AI-driven cyberattacks). The simultaneous advancement of capabilities and security has become the core focus of the next phase of AI model competition.
1. Leading Companies Launch New Models, Restarting the Competition
The AI community has been quite active recently:
- Anthropic has officially released Claude Fable 5.1 and Mythos 5.1, claiming they represent the world's most advanced models for programming and knowledge work.
- OpenAI has announced the upcoming release of its latest model, Astra, but is only testing its advanced cybersecurity features with a limited group of users.
- Google is also making significant progress with its new coding models, aiming to narrow the gap with the other two companies.
The simultaneous release of these models indicates that the competition has intensified. However, the focus is no longer just on which model is the most intelligent; security has become a crucial factor.
2. Anthropic’s New Models: Improved Performance, but Increased Cost per Task
Anthropic’s Fable 5.1 has indeed improved significantly. Third-party evaluations show that it has set new records in various benchmarks, including intelligence scores, programming tests (Terminal-Bench), and financial knowledge tests (τ³-Banking). For example, its intelligence score is 4 points higher than that of its predecessor, and its programming test score of 91.4% is the highest on the platform to date.
However, there are some mixed results regarding pricing:
- The good news is that the cost of cache reads has been reduced by 75%—cache allows the model to reuse previously processed information, saving on computation time.
- The bad news is that the cost per task has increased by 20%. This is because the model’s enhanced reasoning capabilities result in longer outputs, which uses 1.7 times more tokens compared to the previous version, offsetting the cost savings from cache usage.
For corporate users, the frequency at which they can utilize cache (i.e., the cache hit rate) will directly affect their overall costs.
3. OpenAI Places Security at the Forefront: Astra Tested with a Limited Group
OpenAI’s Astra model is unique as it is the first to meet OpenAI’s criteria for key cybersecurity capabilities. It can identify previously unknown security vulnerabilities and automatically develop ways to exploit them without human guidance.
Due to its powerful capabilities, OpenAI is being cautious with its release:
- The model has been training for some time, but its release has been delayed to enhance security measures.
- OpenAI has learned from previous security incidents with Hugging Face and has implemented additional security features, such as rejecting harmful network requests and increasing monitoring to prevent misuse.
- The model is being tested with a small group of users first to ensure its security before expanding the testing scope.
OpenAI CEO Sam Altman emphasized that the simultaneous improvement of capabilities and security is more important than ever.
4. The Entire Industry is Addressing Security Challenges
OpenAI is not the only company focusing on security; the entire industry is taking action:
- Anthropic is analyzing recent security incidents and collaborating with third-party organizations for independent reviews. It has also paused high-risk reinforcement learning training.
- Industry Collaboration: OpenAI, along with Anthropic, Google, Microsoft, and over a hundred other organizations, issued a public letter warning that AI-driven cyberattacks are becoming more frequent and sophisticated. They are calling for collective efforts to strengthen defenses, especially for critical infrastructure such as hospitals, power grids, and water treatment plants, which are facing imminent threats.
Experts from the French AI Security Center believe that these actions are in response to recent security incidents, and companies are finally starting to implement measures to slow down the pace of AI development.
5. Balancing Capabilities and Security: The New Core of Competition
In the past, AI model competition focused on performance. Now, the focus has shifted:
- On one hand, companies are working to improve model capabilities—Anthropic has set new benchmarks in programming, Google is advancing in coding, and OpenAI is making progress with Astra in cybersecurity.
- On the other hand, they are addressing security issues by pausing high-risk training, delaying model releases, and coordinating defenses.
Altman also acknowledged that they previously missed opportunities in the programming domain, but being behind is not a disaster; they can catch up with better models as long as security is not neglected.
In summary, the future competition in AI models will not be about who can run the fastest, but about who can both run fast and stop in time. Both capabilities and security are essential.