Summary of Key Points
China's AI infrastructure is facing a mismatch between "long-term policy planning" and "short-term industry demand," as well as coordination challenges among three critical components: the network infrastructure (the "roads"), computing hardware (the "fuel"), and institutional standards (the "traffic rules"). At the policy level, efforts are made to upgrade the network (to IPv6) and establish rules (for computing power standards) on a five-year basis. However, the industry consumes computing power on a monthly basis and urgently needs to overcome bottlenecks in chip production and scheduling. There is a disparity between surplus computing resources in the western regions and shortages in the eastern regions, with high-end training capabilities being monopolized by giants, making it difficult for small and medium-sized enterprises (SMEs) to break through. To achieve self-sufficiency in AI infrastructure, China must develop its own chip ecosystem and optimize scheduling efficiency.
I. Policy Slows vs. Industry Speeds: The Root of the Mismatch
Policies and the industry operate at different speeds. Policies are typically formulated on a yearly basis (e.g., the five-year IPv6 roadmap or the 12-18-month deployment of AI computing clusters), while the AI industry operates on a monthly scale (companies need computing power for quarterly financial reports, and large models require frequent updates). For example, an AI company may urgently need to expand its computing capacity to support business growth, but building a new center can take 1.5 years; meanwhile, IPv6 upgrades are delayed due to multi-departmental coordination across provinces. This mismatch results in a situation where essential resources are unavailable when needed and redundant resources go unused.
II. Shortcomings in Network Infrastructure, Computing Hardware, and Institutional Standards
The AI infrastructure can be likened to driving a car: the network is the road, computing hardware is the fuel, and institutional standards are the traffic rules. Currently, all three components have issues:
1. Network Infrastructure: The upgrade to IPv6 is underway, but it has not yet reached the level of an intelligent highway that allows for unified computing resource allocation across regions (e.g., using western resources in the east). Many networks still use IPv4, limiting cross-regional communication.
2. Computing Hardware: The problem is not the lack of chips but their quality and supporting components. For instance, the memory bandwidth of domestic Shengteng 910B GPUs (1.6 TB/s) is half that of NVIDIA H200 GPUs (4.8 TB/s), leading to lower training efficiency.
3. Institutional Standards: There are barriers to cross-regional resource allocation, such as data ownership restrictions and conflicts over tax revenue distribution. Additionally, there is a lack of unified standards for comparing the computing power of different chips, causing confusion in transactions.
III. Hierarchical Computing Power Market: Giants Monopolize High-End Resources
The computing power market is segmented into three tiers, making it challenging for SMEs to move up:
1. High-End Training: This requires vast amounts of computing resources (trillions of parameters), which are currently dominated by giants. The cost of high-end clusters is much lower than that of smaller systems.
2. Industry-Specific Tasks: Although the barriers have decreased, costs remain high due to the need for specialized hardware and software adaptations.
3. Inference Tasks: Even with lower entry barriers, high-concurrency, low-latency inference capabilities are in short supply. SMEs often face additional adaptation costs when using cheaper heterogeneous hardware.
IV. Challenges in Implementing Standards
The "Computing Power Standard System Construction Guide" aims to address how to measure and price computing power, but implementation is difficult:
- Technical Barriers: Chip manufacturers resist standardization due to differences in their architectures (e.g., NVIDIA's CUDA ecosystem).
- Interest Conflicts: Local governments and private cloud providers are reluctant to share computing resources for commercial or bureaucratic reasons.
- Evaluation Mechanisms: Current evaluation criteria focus on the number of installed cabinets, not on resource utilization, leading to underutilized centers.
V. The Battle for Efficiency: Moving from Hardware to Efficiency
The future of China's AI infrastructure lies in shifting from a focus on hardware to efficiency. Success will depend on policies that encourage collaboration among governments, enterprises, and technology providers to optimize resource allocation and reduce costs. This includes subsidies for using idle resources, preferential energy policies, and adjustments to evaluation criteria.
Conclusion
The core issue in China's AI infrastructure is the tension between long-term ecological development and short-term commercial interests. To overcome these challenges, a coordinated effort from all parties is needed to shift from building more hardware to improving efficiency. Only by doing so can China gain a competitive edge in the global AI landscape.