虎嗅

Chip Wars Have Moved to a New Arena: Chinese Companies Are Engaging in Token Battles

原文:芯片换了战场,中国公司开打Token仗

Summary of Key Points

This article focuses on Sunrise, an AI chip company incubated by SenseTime, which has seized the opportunity presented by the surge in demand for AI inference. By abandoning the training module and focusing solely on inference-specific GPUs, Sunrise has designed products with large memory capacities, high compatibility, and enhanced hardware-software integration to help customers reduce the cost of tokens (the units used to measure the computational resources consumed by AI). The company offers differentiated solutions tailored to various client needs. Additionally, the excellent synergy between domestic models and chips has enabled Chinese technologies to move from a position of catching up to one of competing on an equal footing in the global market. The article also discusses the competitive landscape, business models, and future prospects within the inference domain.

1. Why Has the Demand for Inference Suddenly Boomed?

In simple terms: The AI “brain” has been developed, and now we need to make it work efficiently on a daily basis.

  • Training vs. Inference: Training is about creating intelligent systems (e.g., training GPT-4 requires significant investment), while inference is about using those systems to perform tasks (e.g., ChatGPT consumes $700,000 per day in inference costs). Previously, most efforts were focused on training; now that AI systems are smart enough, the demand for their practical use has skyrocketed. Deloitte predicts that by 2026, inference will account for two-thirds of the global AI computing power.
  • AI Agents as a Driving Force: Intelligent agents need to plan tasks, utilize tools, and make repeated adjustments (e.g., writing reports requires researching and revising drafts), which consumes significantly more tokens than regular conversations. Agents that monitor systems 24/7 consume even more tokens.
  • Longer Contexts as a Bottleneck: AI systems generating text need to refer to previous information (e.g., articles cannot be written without context). Processing large amounts of context requires substantial memory, which traditional chips struggle to handle, leading to a surge in inference demand.

2. Sunrise’s GPUs: How Do They Make Tokens Cheaper?

Sunrise’s approach focuses on solving three main issues: insufficient memory capacity, slow data transfer, and low efficiency.

  • Large Memory Capacity: Instead of using the high-power HBM memory commonly used in training chips, Sunrise uses LPDDR (low-power) memory from consumer electronics, with capacities up to 600GB—the largest available in China. This allows for more context to be stored and flexible deployment across edge devices and cloud systems.
  • Fast Data Transfer: The latest PCIe Gen6 technology (upgrading from two lanes to four lanes) doubles data transfer speeds. Memory is managed in a layered manner, with frequently used data kept on edge devices and less frequently used data stored in centralized repositories, ensuring smooth resource allocation among multiple users.
  • High Efficiency: By eliminating the training module and dedicating all resources to inference, Sunrise achieves an efficiency of 95% (compared to around 30% for traditional GPUs, or 30% in a factory with only 30 workers).
  • Scalability and Compatibility: 256 chips can be combined to form a “super chip” that can handle the high concurrency demands of large models. The products are 99% compatible with mainstream ecosystems, allowing customers to use them without modifying their code (similar to upgrading a computer without reinstalling software).

3. Who Is Buying Inference GPUs, and How Do They Make Money?

Customers fall into four categories, each with specific pain points that Sunrise addresses precisely:

  • AI Computing Centers: They want to maximize resource utilization and reduce costs per dollar and watt of power consumed.
  • Internet/AI Companies: They need to handle high-concurrency scenarios without lagging and are willing to pay for lower latency.
  • State-Owned and Private Enterprises: They prioritize data security and require local deployment, especially for tasks involving long contexts (e.g., processing complex contracts).
  • Vertical Industries: These industries lack AI expertise and need out-of-the-box solutions (e.g., in manufacturing and healthcare applications).

Sunrise’s profit model relies on reducing customers’ token costs. Revenue comes from both cost coverage and additional profits generated through high-value (e.g., programming, healthcare, with margins over 60%) and low-value (e.g., chat, summarization, with margins below 20%) services, with differentiated pricing strategies.

4. The Opportunities for Domestic Chips: From Catching Up to Competing on an Equal Footing?

Chinese solutions have several compelling advantages:

  • High Compatibility and Low Prices: The cost of tokens for domestic models is one-sixth to one-tenth of those for imported ones. For example, the market value of Zhispu increased by 20 times after its launch, thanks to the combination of domestic models and chips.
  • International Recognition: Companies in Europe, the Middle East, and Silicon Valley are using Chinese open-source models (e.g., DeepSeek) and purchasing Chinese servers and chips because they offer better value and security.
  • Shift in Narrative: The focus has shifted from forcing domestic alternatives to a collaborative approach where both Chinese chips and models can grow together and compete on a global stage.

5. The Future of the Inference Domain: Competition, Trends, and Undervalued Opportunities

  • Competitive Landscape: There are several types of players: overseas giants (expensive and with unstable supply), companies that integrate training and inference (but not specialized in inference), ASICs (specialized but obsolete when models evolve), and Sunrise (specialized in inference with general compatibility).
  • Future Trends: Tokens are expected to become cheaper, similar to how data traffic has decreased in cost. However, high-value tokens (used in AI agents and healthcare applications) will likely see price increases due to rapid demand growth. Lower costs will enable more applications to be developed, further driving overall demand.
  • Undervalued Opportunities:
  • Memory: Memory costs account for a significant portion of inference costs; large-capacity, low-cost memory solutions are currently underutilized.
  • Energy Efficiency: Data centers, which convert electricity into tokens, can benefit from energy-saving technologies (e.g., liquid cooling and efficient power supply).
  • High-Quality Data: As AI integrates with the physical world (e.g., autonomous driving), there will be a growing demand for high-quality data, leading to the emergence of leading data companies.

In essence, this article highlights that the AI industry is transitioning from a phase of expending resources on model development to one of optimizing performance and efficiency. Inference chips are at the heart of this transition. Companies that can make tokens more affordable and reliable will have a significant advantage. Domestic chip manufacturers are emerging as strong contenders in this field, proving their competitiveness.