Summary of Key Points
In the first half of 2026, the widespread adoption of AI Agents led to the largest architectural upgrade in the cloud industry in two decades. The cloud has evolved from providing "Internet-based cloud" services for applications and "GPU cloud" support for model training to a "Agent cloud" that enables AI to actually perform tasks. The fundamental focus of the cloud has shifted from simply providing quick answers to delivering high-quality task completion. This shift has sparked a surge in demand for components such as GPUs, CPUs, HBM (High-Bandwidth Memory), and optical modules along the supply chain. New cloud vendors (like NeoCloud) have emerged rapidly, while traditional cloud giants are reorganizing their capabilities. The competitive focus has moved from selling computing power to ensuring the stable operation and scaled delivery of AI Agents. The measure of value has also changed from "Token efficiency" (cost per speed) to "intelligence output per Token" (achieving more tasks with fewer resources).
The Three Transformations of the Cloud: From "Renting Servers" to "Enabling AI to Work"
The evolution of the cloud can be divided into three stages, each corresponding to different needs:
1. Internet-based cloud (since 2006): Computing power was purchased in a similar manner to electricity. Early companies bought their own servers and set up data centers, which either resulted in insufficient capacity (leading to system crashes during business surges) or wasted resources (when devices were not utilized). Amazon EC2 made computing resources available on a pay-as-you-go basis, solving this issue. At this stage, cloud services primarily supported applications like websites and apps, with the goal of making software run more efficiently and cost-effectively.
2. GPU cloud (after ChatGPT in 2022): Clouds began to support model training. Large models require massive parallel computing, and GPUs, which are better at sequential tasks (not suitable for parallel processing), became a core component. Cloud vendors started offering AI platforms (such as Baidu's Qianfan and Alibaba's Bailian) that encapsulated model training and inference services, eliminating the need for companies to build their own infrastructure.
3. Agent cloud (since the second half of 2025): The focus is on enabling AI to perform actual tasks. Agents are no longer just answering questions but executing tasks, such as cooking with robots or automating business processes. This requires clouds to provide more advanced capabilities, including multi-round reasoning, tool invocation, memory management, and security isolation—essentially, providing an "agent manager" (Agent Harness) for the AI.
The New Infrastructure of the Agent Cloud: More Than Just Computing Power
For agents to function reliably, a new set of components is needed, forming the foundation of the cloud:
- What is an Agent Harness? It acts as the "operating system" for agents, working between the model and applications. Its responsibilities include:
- Task scheduling (e.g., deciding which tool to use first when cooking with a robot).
- Tool invocation (e.g., checking the weather, using API interfaces, controlling hardware).
- Memory management (e.g., remembering user preferences, such as not liking spicy food).
- Security isolation (e.g., preventing agents from accessing internal company data).
- The Bucket Effect: The overall performance of an agent is determined by the weakest link in its infrastructure. For example, if an agent mistakenly assumes the existence of a non-existent tool or performs tasks too slowly, it can fail. Therefore, cloud vendors must optimize every component.
The Flourishing Supply Chain: Which Hardware Components Are in High Demand?
The emergence of the Agent cloud has changed the demand for hardware from individual components to integrated combinations:
1. GPU + CPU: GPUs handle model inference, while CPUs are used for task scheduling and tool invocation (e.g., coordinating database queries). As a result, CPUs, which were once marginalized, have seen a resurgence.
2. HBM + Optical Modules: Agents require fast storage and high-speed data transfer between GPUs. HBM and optical modules have become scarce, leading to significant price increases for related companies.
3. Rise of NeoCloud Vendors: New cloud vendors like CoreWeave and Nebius abroad offer customized computing solutions, taking over orders from traditional giants (e.g., Nebius' stock price increased by 10 times, and it was included in the NASDAQ 100 index). In China, traditional cloud providers (such as Baidu Smart Cloud and Alibaba Cloud) have upgraded their services to support AI natively, without a separate NeoCloud market.
The New Arena for Cloud Vendors: From Selling Computing Power to Providing Comprehensive Solutions
The competitive focus of cloud vendors has shifted:
- Microsoft: At Build 2026, Microsoft proposed that Agents will replace apps and operating systems, introducing a full-stack Agent service covering perception to distribution.
- Google: Google discontinued Vertex AI after five years and replaced it with the Gemini Enterprise Agent Platform, along with its eighth-generation TPU optimized for agents.
- Baidu Smart Cloud: Baidu is developing "Huisi Kaiwu" embodied intelligent systems, emphasizing the importance of agent success rate (e.g., robots should not need to try multiple times to perform a task correctly).
The Change in Value: From Token Efficiency to Intelligent Use of Tokens
Previously, cloud vendors competed on how efficiently they could generate tokens (e.g., cost per token or energy efficiency). In the era of agents, the focus is on the intelligence delivered per token—using the same amount of resources to complete more complex tasks (e.g., generating reports automatically, analyzing data, and providing recommendations).
This requires cloud vendors to offer a "full-stack service" that covers all aspects, from underlying computing power to the agent management layer and final implementation in business scenarios. For example, companies can rent GPUs for model training, use the agent platform to develop intelligent systems, and then purchase ready-made agent applications, gradually upgrading their infrastructure.
In summary, the advent of the Agent cloud has transformed the cloud from a provider of computing power into a service provider of AI productivity. In the future, those who can successfully address the challenges of deploying agents—stability, security, and efficiency—will emerge as winners in the AI cloud market.