Summary of Key Points
This interview focuses on the transition of AI from the “Model Era” to the “Agent Era,” with a central discussion on Agent Infra (the infrastructure for intelligent agents), a critical area where giants are making stealthy strategic moves. Agents are capable of performing real tasks, but their large-scale deployment in production still faces bottlenecks due to insufficient model capabilities. Agent Infra has evolved from being purely software-based systems to complex “hardware-software integrated” systems, with GPUs at the core. The interview compares the different approaches taken by domestic and international giants (Google, NVIDIA, OpenAI, Anthropic, etc.) in developing Agent Infra, analyzes the three major technological challenges of computing power, storage, and networking, and highlights the unique path taken by domestic infrastructures due to their ability to adapt to multiple chips and open-source advantages. It concludes by emphasizing that this is a golden age for engineers, as there is now greater potential for exploration and impact in this field.
1. The Transition from Pure Software to Hardware-Software Integration in the Agent Era
The AI infrastructure of the past was purely software-based—similar to the CPU era, where chips were general-purpose and software could run directly on them. However, with the advent of the Agent Era, the infrastructure has changed:
- GPUs have become the core: GPUs offer high computing power but are not versatile; they are specialized for specific tasks and cannot perform a wide range of functions like CPUs. Therefore, the infrastructure must integrate hardware (GPU chips) with software (model optimization and tool invocation) to create a “hardware-software integrated” system.
- Changing Tasks: In the past, infrastructure was responsible for simply invoking models (e.g., using ChatGPT for conversations). Now, agents need to complete real tasks (such as writing code or handling files), which requires a comprehensive environment that includes tools, memory management, and process control. This means the infrastructure must support multiple iterations and stable operations.
2. Agents Can Create Demos, but Difficulties in Production: Insufficient Model Capabilities Are the Key Issue
It’s easy for agents to create small demos (e.g., writing a simple report), but deploying them in production environments (e.g., replacing backend engineers) is challenging due to limitations in model capabilities:
- Limited Coverage of Boundary Conditions: Production tasks involve handling various unexpected situations, and current models are not yet capable of addressing all of these.
- Inability to Handle Position-Level Tasks: Agents can perform single tasks (e.g., data retrieval), but to replace a full position (e.g., a product manager), they need to integrate multiple skills and processes, which current models do not yet possess.
- Unreproducible Results: Production environments require consistent results for each execution, but agents often provide different outcomes for the same issue, indicating a need for infrastructure solutions to ensure predictability.
3. The Giants’ Competition in Agent Infra: Different Approaches
Giants are competing for control of Agent Infra, which is akin to the future “operating system”:
- Google: Offers a full-stack integration (self-developed TPU chips + Gemini models + JAX software), with deep expertise. They have been developing TPU chips for 10 years and have achieved perfect compatibility with models and software, resulting in low costs and high performance.
- NVIDIA: Focuses on vertical integration from chips to systems. They not only sell GPUs but also develop the CUDA software ecosystem and hyper-node interconnection technologies (enabling high-speed collaboration between multiple GPUs). They even acquired Groq to develop advanced inference chips, creating a strong ecological barrier.
- OpenAI: Adopts an ecosystem platform approach, aiming to create an “AI App Store” (GPT Store). They collaborate with Broadcom for chip development to ensure sufficient computing power and use their ecosystem to retain users.
- Anthropic: Emphasizes lightweight solutions for high-value tasks. They do not develop their own chips but partner with large companies for computing resources, focusing on developing model operating systems (e.g., Claude Code) to assist users in performing valuable tasks like coding.
4. Overcoming Three Major Technical Challenges: Computing Power, Storage, and Networking
The Agent Era presents three key infrastructure challenges:
1. Computing Power: Agents need to invoke numerous tools and process large amounts of context, requiring more powerful inference clusters. They also need to be adapted to domestic chips (given the unique domestic market conditions), which requires addressing performance differences between different chip types.
2. Storage: Agents require significant amounts of storage for context and memory; however, high-end storage (e.g., HBM) is in short supply. Storage capacity expansion is slow (taking more than two years to build new production lines), and the surge in AI demand has led to increased storage prices.
3. Networking: Agent collaboration requires fast interconnection. Giants are developing “hyper-node” technologies that connect multiple GPUs or even entire cabinets using high-speed buses, increasing communication bandwidth by more than ten times.
There is also an emerging trend: Increased CPU Usage: Agents invoke tools more frequently than humans, leading to higher CPU consumption. If CPU usage rises, it indicates that agents are truly being deployed in production environments.
5. The Unique Path of Domestic Infra: Adaptation to Multiple Chips and Open Source
Domestic infrastructures differ from those abroad and have their own strategies:
- Adaptation to Multiple Chips: While foreign companies mainly use NVIDIA chips, domestic firms (such as Huawei and Baidu with Kunlun Chip) need to support domestic chips. This has led to the development of unique technical approaches.
- Open-Source Advantages: Domestic open-source models (e.g., Wenxin Yiyan, Llama series) are more active, allowing for collaborative optimization of architectures and faster infrastructure iteration.
- New Metrics: Baidu has introduced “Daily Active Agents” as a new metric to replace the traditional “token consumption.” This measure focuses on the actual value of tasks completed by agents (e.g., the number of agents performing high-value work daily).
In conclusion, the interview highlights that this is a great era for engineers, as there are numerous opportunities for innovation in Agent Infra, from chips to software, and from models to tools. Engineering skills are becoming increasingly important. The success of AI systems depends not on fleeting algorithmic breakthroughs but on the ability to build and optimize entire systems effectively.
This interview underscores that while agents represent the future of technology, the underlying infrastructure is crucial for their practical application. Giants are competing in this invisible battlefield, and engineers have the opportunity to create solutions that can transform the industry.