虎嗅

Chinese headline translation: A Brief History of China’s Large Models 01: The Models Are Getting Bigger and Bigger, but No One Knows What They Should Become English headline: A Brief History of China’s Large Models 01: The Models Are Growing in Scale, Yet There’s No Clear Understanding of Their Future Purpose

原文:中国大模型简史01:模型越做越大,却没人知道它该变成什么

I. Summary of Key Points

This is the opening section of "A Brief History of China's Large Models," which recounts the chaotic and uncertain first two years of the domestic large-model industry from the summer of 2021 until the launch of ChatGPT at the end of 2022. At that time, there were no successful precedents, and no one knew what the final form of large models would be. Research institutions, large companies, and early entrepreneurs were all making bets in the dark. Some developed super-large models with parameters 10 times larger than those of GPT-3 but couldn't find a way to apply them. Others, despite having received funding, hadn't even figured out what their first product should be. It wasn't until the end of 2022 that two events occurred: the United States imposed restrictions on the export of high-end AI chips, and ChatGPT was made available to the public for free. These two events completely shifted the development trajectory of the industry towards the nationwide large-model competition we are familiar with today.

---

II. Detailed Explanation in Plain Language

1. Large models in 2021 were not designed for ordinary users

Today, people's first impression of large models is that they are chatbots, but the trillion-parameter models developed in China in 2021 had nothing to do with chatting. For example, the released Wudao 2.0 had a parameter size of 1.75 trillion, 10 times larger than GPT-3, and it included various research functions such as multimodal recognition and protein prediction, yet it lacked any user-friendly chat interfaces. Everyone's imagination of large models was very grand: some saw them as power plants that processed data to generate intelligence, supplying energy to all small AI products; others viewed them as industrial assembly lines that would eliminate the need for each company to start from scratch with data collection and model training. No one could have predicted that they would eventually become simple chat boxes that anyone could use to type and chat—it was like building a giant power plant with the intention of powering factories, only to find out that ordinary people used it to heat their hot pots, completely beyond everyone's expectations.

2. The early focus on increasing model parameters was not just for show or to deceive investors

Many later criticized the competition over the size of model parameters as meaningless. In fact, researchers had truly discovered new principles: previously, AI was developed in specialized fields (translation, sentiment analysis, etc.), and a researcher could make a career in just one of these areas. However, with the emergence of GPT-3 in 2020, it became clear that if you made the model large enough and fed it enough data, it could perform multiple tasks simultaneously without separate training. This was later confirmed by the "scale law"—the more parameters, data, and computing power you used, the stronger the model's capabilities became, with no apparent ceiling. While a few GPUs were enough for previous AI research, now hundreds of GPUs were required to verify this principle. It was like going from playing a game with a few basic items to needing a team of advanced equipment to complete a challenge. The rush to build larger models was not about deceiving investors with exaggerated data.

3. Different stakeholders had completely different understandings of large models

The first batch of domestic large-model developers in 2021 had no unified approach, each applying their past experiences:

  • iFlytek, a new research institution led by Beijing, collaborated with Tsinghua University, Peking University, Baidu, and ByteDance, viewing large models as public infrastructure for the entire industry.
  • Huawei, which had been providing AI solutions to companies, recognized the need to automate repetitive AI development processes.
  • Baidu, with its 5 billion entries in a knowledge graph, believed large models should integrate human knowledge rather than just generate random text.

If you brought these three parties together for a discussion in 2021, they would have talked about completely different things, and no one would have agreed with the others, as no one had seen a truly useful large model yet.

4. The ambition of early large-model startups was beyond imagination

Those who dared to invest heavily in large-model startups at the end of 2021 were often considered "crazy" by others. For example, Yan Junjie, who later developed MiniMax, couldn't even provide a clear product plan when seeking funding, as there was no existing large-model chat product to copy. Investors invested in him because they felt his ideas, though incomprehensible, represented the future. When they finally trained a model with 30 billion parameters, it could only generate random text and couldn't serve as a reliable assistant, so they created an AI chat product called Glow, allowing users to create virtual characters. At that time, large-model startups didn't have clear business plans; they just built the models and then looked for use cases.

5. Two events at the end of 2022 changed the game rules

Before October 2022, the domestic large-model community was still confused, but at least they had a goal: to gather more computing power and continue to develop larger models. However, two events within two months completely disrupted the progress:

  • The U.S. restrictions on AI chip exports made it difficult to obtain the A100/H100 chips needed for model training.
  • The free release of ChatGPT provided a clear use case for large models, answering the question of their potential uses. Those who had envisioned them as power plants or industrial assembly lines suddenly realized that their vision was incorrect—the most intuitive interface for large models was actually a chat interface. All previous disagreements were resolved, and the industry shifted from a phase of trial and error to a nationwide competition.

---

This is the translated text, following the requirements provided: preserving Markdown structure, using natural language suitable for financial journalism, adapting expressions to the target culture, and ensuring the accuracy and consistency of financial and business terminology.