虎嗅

OpenAI's First Self-Developed Chip, "Chili," Has Been Released: Behind the 9-Month Breakthrough at the Speed of Light, Something NVIDIA Fears Most Is Happening

原文:OpenAI首颗自研芯片“辣椒”问世:9个月光速成功背后,英伟达最怕的事正在发生

Summary of Key Points

OpenAI has unveiled its first self-developed intelligent processor, Jalapeño (named after a Mexican chili), which is an ASIC chip specifically designed for large-scale model inference (not a general-purpose GPU/TPU). The development and manufacturing of the chip took only 9 months, claiming to be the fastest semiconductor development cycle in history. AI models were also involved in the optimization of the chip design. Partners include Broadcom (responsible for chip implementation and mass production) and Celestica (for assembly), with the strategic goal of creating a “full-stack closed loop” where OpenAI controls every aspect from model development to chip production and deployment. Currently, the chip is only used for inference tasks (not training, which still relies on NVIDIA GPUs). The aim is to reduce inference costs by 50% compared to typical AI GPUs, with plans for multiple generations of chips to ultimately achieve a computing power output of 10 GW.

Detailed Analysis

1. How Did OpenAI Achieve Such Rapid Development in 9 Months?

Traditional chip development usually takes 18-36 months, but OpenAI managed to compress it down to 9 months thanks to three key factors:

  • Close collaboration between software and hardware teams: OpenAI engineers worked alongside Broadcom from the very beginning, adopting a iterative design process that allowed for continuous adjustments and doubled efficiency.
  • AI as a design assistant: The most time-consuming part of chip design is testing thousands of iterations. AI can analyze historical data, generate code, help identify bugs, and optimize the layout, taking on the most time-consuming tasks.
  • Customized design rather than modification: Instead of modifying existing chips, the new chip was designed from scratch to meet OpenAI’s specific requirements for cores, memory, and service models. OpenAI even brought in Richard Ho, a core engineer from Google TPU, to apply their experience in AI-assisted chip design.

This speed suggests either that OpenAI had already made significant preparations or that AI significantly enhances chip design efficiency beyond expectations.

2. Why Focus on Inference Instead of Training?

Inference and training are two stages of AI development:

  • Training: Similar to teaching a student, this involves feeding large amounts of data into a model to develop its capabilities (high initial cost and technical barriers).
  • Inference: Like a student answering questions, it’s the continuous process of processing user queries (lower ongoing cost, but still more expensive than training).

OpenAI chose to focus on inference for practical reasons:

  • Inference is the current bottleneck: The main issue with inference is data transfer—the time it takes to move model weights from memory to computing units. Jalapeño is optimized to reduce this process and improve efficiency.
  • Cost savings: A 50% reduction in costs could be substantial, especially for OpenAI, which spends billions of dollars on computing power annually.
  • Avoiding direct competition with NVIDIA: Training chips are NVIDIA’s core domain; targeting the inference market first reduces risks and allows for faster adoption.

Once inference is established, OpenAI can later tackle training.

3. The Full-Stack Closed Loop: What NVIDIA Worries About

NVIDIA’s major customers are beginning to switch: Google has TPU, Amazon has Inferentia, Microsoft has Azure Maia, and now OpenAI has its own chip. This shift is crucial because computing power is the lifeblood of AI, and no company wants to be dependent on others.

OpenAI’s full-stack closed loop means:

  • Reducing reliance on Microsoft and NVIDIA: Previously, both inference and training relied on NVIDIA GPUs through Azure. With its own chips, OpenAI can reduce this dependency.
  • Gaining pricing and control: By designing its own chips, OpenAI can better match model requirements and avoid being constrained by suppliers’ technology roadmaps.
  • Ambitious goals: OpenAI aims to provide 10 GW of computing power by 2029 (equivalent to the output of 10 nuclear reactors). Broadcom also notes that this is just the beginning of a multi-generation roadmap, with plans for deploying gigawatt-level data centers.

NVIDIA’s customers becoming competitors is certainly concerning for them.

4. What Benefits Does This Chip Bring to Ordinary Users?

Although the chip will not be deployed until the end of the year, the long-term implications are significant:

  • Cheaper AI services: Reduced inference costs could lead to lower fees for AI services like ChatGPT, making advanced AI more accessible to students, small businesses, and individual developers.
  • Faster response times: Optimized data transfer will result in faster responses and a smoother user experience.
  • Accelerated AI adoption: Lower inference costs will enable more industries (education, healthcare, small businesses) to adopt AI tools, such as using AI for customer service and content creation.

However, the actual performance of Jalapeño remains to be seen. The trend is clear: AI companies are taking control of their computing power, a trend that is unstoppable.

5. What’s Next for OpenAI?

OpenAI’s strategy aims to create a self-sustaining “computing power flywheel”:

  • Designing chips to reduce inference costs.
  • Making AI more accessible to everyone, generating more data.
  • Training stronger models.
  • Continuously optimizing chips for even lower costs and better performance.

The timeline is as follows:

  • Jalapeño will be deployed by the end of 2026.
  • The next generation of chips will be launched in 2028.
  • Annual iterations will follow.
  • 10 GW of computing power will be achieved by 2029.

If this flywheel takes hold, OpenAI will not only reduce its dependence on suppliers but also build a stronger competitive advantage in the AI industry. After all, those with more efficient computing power will have a significant edge.

In summary, Jalapeño is more than just a chip; it represents a critical step for OpenAI in taking control of the entire AI ecosystem. This marks the beginning of a race to dominate the production and distribution of computing power.