Summary of Key Points
In June 2026, OpenAI collaborated with Broadcom to release the Jalapeño, the first custom AI inference chip designed specifically for large language models (LLMs) to answer questions and generate content. This chip features a rapid development cycle (from design to manufacturing in just 9 months) and high efficiency (outperforming current leading technologies in terms of performance per watt). It also utilizes AI-driven design methods and optimized resource allocation to maximize its effectiveness. Behind this initiative is OpenAI's strategy to reduce operational costs, reduce reliance on a single chip supplier (such as NVIDIA), and stay ahead of industry competitors (e.g., Google, which has its own TPU chips). The Jalapeño is planned for initial deployment by the end of 2026, with further iterations expected in the future.
Detailed Analysis
1. What exactly is this chip?
The Jalapeño is an inference chip, designed to process user requests after AI models have been trained (for example, when you ask questions to ChatGPT). Unlike general-purpose GPUs (such as NVIDIA's A100), it is an ASIC (Application-Specific Integrated Circuit), meaning it is tailored for inference tasks only and is more efficient than general-purpose chips. Its advantages are clear:
- Energy-efficient and powerful: Early tests show that it performs better per watt of power consumption compared to the best current chips.
- Near-optimal efficiency: By minimizing data movement within the chip, it balances computational, memory, and networking resources, bringing the chip's actual performance close to its theoretical maximum. This is a significant improvement, as many chips have strong theoretical capabilities but suffer from inefficiencies in practice.
- Capable of running advanced models: Engineering prototypes have already been used to run large models like GPT5.3 and Codex at production-grade frequencies and power levels.
2. Why does OpenAI want to develop its own chip?
OpenAI previously relied on renting NVIDIA GPUs from Microsoft, but now it needs to take control due to several factors:
- High costs: The computational resources required for inference (answering questions) have surpassed those needed for training models, making GPU rentals increasingly expensive.
- Supply chain risks: Relying solely on NVIDIA exposes OpenAI to supply disruptions or price increases, which can be detrimental.
- Competitive pressures: Competitors like Google have their own TPU chips that are cost-effective and well-integrated with their models, offering significant profit margins. OpenAI must address this by developing its own chips to maintain a competitive edge.
3. The collaboration model
This project is a collaborative effort among three parties:
- OpenAI: Responsible for the underlying architecture design, determining which parts of the chip will handle computation and data storage.
- Broadcom: Converts the architectural design into a physical silicon chip and provides networking hardware (e.g., the Tomahawk chip for high-speed communication between chips).
- Celestica: A Canadian electronics manufacturer that assembles the chip onto circuit boards and integrates it into rack systems suitable for data centers.
This collaborative approach has been highly efficient, resulting in a prototype within 9 months, with OpenAI's own AI models being used to accelerate the design process.
4. Strategic significance
Developing its own chips marks a shift for OpenAI from being a “software company” that relies on external hardware to becoming a “full-stack player”:
- Greater control: OpenAI can optimize all aspects of its technology, from algorithms to hardware, without being constrained by suppliers.
- Cost savings and efficiency improvements: Custom chips are better optimized for OpenAI’s models, potentially reducing computational costs significantly (for example, Google has seen substantial profit increases since adopting TPU chips).
- Future vision: This move is part of OpenAI's long-term strategy to build a comprehensive infrastructure. Future generations of chips will support larger-scale AI applications, such as gigawatt-level data centers capable of handling massive amounts of computational power.
5. What’s next?
The Jalapeño is scheduled for initial deployment by the end of 2026, with additional iterations planned. This development could:
- Inspire others: Encourage other AI companies to develop their own chips or collaborate more closely with hardware manufacturers to make AI computing more affordable and efficient.
In summary
OpenAI’s decision to develop its own chips is not a whim but a strategic move to avoid being bottlenecked by external factors, reduce costs, and improve efficiency. By taking this step, OpenAI aims to gain a competitive advantage in the emerging “computational economy.”