Summary of Key Points
Currently, tech giants are all betting on “full-stack AI” – from chips and computing power to models, applications, and endpoints – in an attempt to control every link in the AI industry chain. This impulse stems from two main motivations: firstly, the fear of being constrained by others (for example, model companies lacking computing power, or cloud providers fearing the loss of customers); and secondly, the successful experience of “borderless expansion” in the internet era, which makes them believe they can do anything. However, full-stack expansion comes at a heavy cost: it consumes excessive amounts of capital, leading to negative free cash flows, increased internal resource competition and coordination costs, and even slowing down their core businesses. In the long run, full-stack strategies are more akin to a manifestation of anxiety in an immature industrial division of labor. As external transaction costs decrease (for instance, APIs and standard interfaces make it easier for different companies’ products to connect), these giants will likely return to a more efficient division of labor where each focuses on its own expertise.
Why Do Giants Want to Dominate AI? Anxiety Is at Work
The drive for full-stack AI by these giants is essentially a fear of losing out. On one hand, every link in the AI chain could become a new source of competitive advantage: chips (like NVIDIA’s), models (like Anthropic’s), and endpoints (like Apple’s distribution rights). Any absence in these areas could result in losing market share to competitors. For example, model companies that rely on others’ GPUs may face price increases or supply disruptions; cloud providers that only offer computing power could see customers poached by model companies.
On the other hand, the success of the internet era has led them to believe that expansion is the key to everything. Examples like Zhang Yiming’s concept of a “Super Company” and Wang Xing’s approach to “borderless expansion” demonstrate their willingness to invest heavily in opportunities. They think that small features today could grow into large platforms tomorrow, and thus they cannot afford to miss any potential opportunities. This mindset has turned full-stack strategies from a technical choice into a strategic anxiety.
There Are Two Types of Full-Stack Expansion: One Is Useful, the Other Is Greed
Full-stack expansion is not monolithic; it can be divided into two categories:
1. Vertical Integration (expanding around core capabilities): For example, model companies moving into chip manufacturing and data center operations (OpenAI’s Jalapeño chip) or cloud providers developing their own models (Microsoft’s MAI model). This reduces external dependencies; by producing their own chips, they don’t have to rely on NVIDIA, and by developing their own models, they retain customer loyalty. Such expansion is valuable as it improves synergy between upstream and downstream processes (for instance, Microsoft’s custom chips increased Copilot’s inference speed by 40%).
2. Horizontal Expansion (trying everything): After developing a model, companies may also venture into chatbots, coding tools, browsers, and hardware endpoints. This is not about addressing weaknesses but about not wanting to miss any potential entry points. For example, developing agents could prevent customers from being poached by competitors, and creating endpoints ensures control over distribution. However, this type of expansion disperses resources, dividing computing power, talent, and funds across multiple businesses, which can slow down core development.
The High Cost of Full-Stack Expansion
The consequences of full-stack expansion are evident:
1. Excessive Capital Consumption Leading to Negative Cash Flows: This year, companies like Alphabet and Microsoft are expected to invest $650 billion in AI infrastructure, a nearly 60% increase from last year. Alibaba’s capital expenditure in the second quarter rose by 75%, and Tencent spent $8 billion, both experiencing negative free cash flows. ByteDance even took out a $20 billion offshore loan to boost its AI efforts. While these investments pay off in the long term, the immediate pressure on cash flow is immense.
2. More Difficult Internal Coordination Than External Cooperation: Google DeepMind’s reorganization illustrates this issue: the Gemini model team and the cloud division competed for TPU resources, with conflicting goals (long-term AGI research versus rapid product deployment) leading to inadequate investment in key areas. Alibaba’s Qwen team also reported difficulties in accessing resources within the company, despite the ease with which external companies use its cloud services. The larger the organization, the higher the coordination costs, making external division of labor more efficient.
Full-Stack Is Not the Endgame; Division of Labor Is the Way Forward
Economist Ronald Coase argued that companies exist because market transaction costs are high (e.g., negotiating with suppliers is time-consuming). However, if internal management costs are even higher, it’s better to outsource certain tasks. This is precisely what is happening in the AI industry:
1. Rising Internal Coordination Costs: As OpenAI has evolved from a research lab to a product company and then to an infrastructure provider, its organization has become more complex, requiring professional managers and traditional management structures, which reduces flexibility.
2. Declining External Transaction Costs: APIs make model capabilities readily available, and standards like MCP and A2A enable seamless integration of agents and tools from different companies. For example, Anthropic’s MCP is used by ChatGPT and Gemini, and Google’s A2A interfaces are open-sourced, making cross-company collaboration more cost-effective than developing everything in-house.
In the future, the AI industry will likely return to a more efficient division of labor, with chip companies focusing on chips, model companies on models, and application companies on their respective areas, all connected through standard interfaces. Full-stack strategies will be a temporary phase.
Are Giants Starting to Realize This? The Shift at OpenAI and DeepMind
Recent developments indicate that these giants are adjusting their strategies:
- OpenAI: Altman has clearly stated that they cannot try to do everything and will focus on building a platform that provides user access and APIs for developers, cutting back on less successful products like Sora and Atlas to concentrate resources on core areas.
- DeepMind: Hassabis has stepped back from daily management to focus on AGI research, with operations handed over to more experienced commercializers, and some teams have been reintegrated into Google’s structure. These changes show that giants are realizing that full-stack expansion is not a guaranteed path to success; it’s better to deepen their core competencies.
In Conclusion
Full-stack AI represents a collective anxiety, but it will eventually return to rationality. Only those who can make their core capabilities irreplaceable will establish a solid foothold in the AI era, rather than trying to dominate everything.