Summary of Key Points
The issue in China's computing power landscape is not a general surplus, but rather a mismatch in structure: on one hand, there is a shortage of high-quality computing power required for tasks such as large model training and AI research (e.g., high-bandwidth clusters and mature software ecosystems); on the other hand, many existing computing resources remain idle due to inadequate compatibility, scheduling difficulties, and a lack of customers. Although more than 60% of the country's computing power has been integrated into a unified monitoring system, a viable end-to-end delivery mechanism that converts physical computing power into services that enterprises can utilize, afford, and use to complete tasks has not yet been established. The key to future competition will not be who can access the most computing power, but who can effectively bridge the gap between available resources and actual workloads.
Detailed Analysis
1. The nature of the computing power shortage has changed: from "availability" to "usability"
In the past, the shortage referred to a lack of hardware; now, it means having the equipment but being unable to use it or failing to find the right equipment. For example, companies that want to train large models need high-speed interconnected clusters composed of hundreds of identical GPUs, but such clusters are either unavailable or their software environments are incompatible. In some areas, intelligent computing centers with many GPUs have only a 30% utilization rate because they lack suitable tasks. This situation of simultaneous shortages and idleness reflects a mismatch between supply and demand—there is a shortage of high-quality, functional computing power, while there is idle capacity that cannot be effectively utilized.
2. Monitoring does not equate to usability: five hurdles must be overcome to make computing power usable
Even though over 60% of the country's computing power is being monitored, this only provides information about where it is located and what type of chips are used, but it is still far from making it available for practical use by enterprises. The process from having computing power to using it for payment involves five stages of verification:
- Equipment deployment: Has the hardware been properly installed and made available? (According to the MIIT, the overall deployment rate is 71.4%, meaning nearly 30% of the equipment has not yet been deployed.)
- Platform monitoring: Can the computing power be accessed through a unified platform? (Currently, 60% of the computing power is visible, but the total number of resources is still increasing, limiting the significance of this metric.)
- Interface compatibility: Can enterprises connect to this computing power via APIs? (Many resources have closed or inconsistent interfaces.)
- Task compatibility: Can the computing power run the required models? (Different chips and software stacks may require different adaptations; for example, a model that runs on one chip might not work on another.)
- Customer willingness to pay: Are enterprises willing to pay for the service? (If the cost is too high or the reliability is poor, they will not use it.)
At each stage, some of the computing power will be eliminated, so out of 100P of nominal computing power, only about 30P might actually generate revenue.
3. The challenges of scheduling: incompatibility between tasks and resources
Even if idle computing power is identified, it may not be allocated to the right enterprises due to various constraints:
- Task requirements: Large model training, for instance, requires high-bandwidth, low-latency clusters. Idle GPUs in different locations might not be able to be combined into a usable cluster.
- Resource heterogeneity: Different chips (e.g., GPUs and NPU) have different operating systems, requiring costly re-adaptations of models.
- Data compliance: Medical and financial data cannot be freely transferred across regions and must be processed locally, limiting the use of idle resources in other areas.
- Cost considerations: Although computing power in the western regions is cheaper, the cost of data transfer and additional software may exceed the savings from using cheaper electricity.
These constraints mean that the amount of idle computing power available nationwide does not necessarily equate to the amount that enterprises can actually utilize.
4. The limitations of the national computing power network
While the network can help alleviate the mismatch, it cannot eliminate all differences:
- Suitable tasks: Some tasks (e.g., film rendering and scientific simulations) that are not time-sensitive can be scheduled to use idle resources as long as the software is compatible and the data can be transferred.
- Incompatible tasks: Highly dependent tasks (e.g., large model pre-training) require all available chips to work together, and only local high-quality clusters can handle them. Therefore, even with a improved network, the issue of simultaneous shortages and idleness will not completely disappear.
5. The next round of competition: shifting from hardware to delivery
In the past, the focus was on building larger computing centers and accessing more power; now, the focus is on turning physical computing power into deliverable services. The core competencies in the future will include:
- Hardware manufacturers: Can they provide stable clusters and mature software ecosystems that enable models to run seamlessly?
- Scheduling platforms: Can they quickly match tasks with appropriate resources and address compatibility, networking, and settlement issues?
- Service providers: Can they offer reliable networks (e.g., with guaranteed latency and fast fault recovery)?
Those who can package these services into marketable offerings (e.g., "100P of computing power for 3 days of training, refund if failed") will gain a competitive advantage in the future computing power market. Currently, there is a shortage of providers who can integrate all these elements effectively, as customers often have to manage each step themselves, leading to higher costs and greater risks.
Conclusion
China's computing power infrastructure construction is not yet complete, but the focus has shifted from building more capacity to ensuring that it is actually usable. The winner will be those who can transform static hardware into dynamic services that enterprises can rely on to complete their tasks on time and with quality.