虎嗅

"The Base Model Is Dead."

原文:The Base Model Is Dead。

Summary of Key Points

This article discusses a significant shift in the training logic of large models: In the past, the focus was on the Base Model, which was designed to absorb as much human knowledge as possible. However, the role of the Base Model has changed. It is no longer merely a “container” for knowledge but has evolved into a “toolbox of basic skills,” providing fundamental capabilities such as Python and code structures. The ultimate complex abilities of the model—such as writing software and solving complex tasks—are now achieved through Reinforcement Learning (RL) and a new approach called Mid-training. Additionally, the pre-trained data has shifted from being a “hodgepodge of various web pages” to being more targeted and precise. The boundaries between different training phases have become increasingly blurred, simplifying the overall process into two main steps: supervised learning for foundational skills + RL for application development.

1. Base Model: From a “Super Library” to a “Toolbox”

In the past, the Base Model was like a super library that incorporated information from the internet, Wikipedia, and books. During the GPT-3 era, 85% of the pre-trained data came from web pages and Wikipedia, and it was believed that the better the pre-training, the stronger the model would be. Post-training involved only fine-tuning the model’s chat behavior, with RL playing a limited role.

Now, the goal of the Base Model is different. It no longer needs to directly perform complex tasks (such as writing 10 hours of software code) but must master fundamental skills like Python syntax, code structure, Git operations, and API calls. This is similar to learning how to cook; the Base Model teaches you the basic steps (chopping vegetables, flipping a pan, mixing sauces), while RL helps you combine these steps to create a complete dish. The Base Model has evolved from being “all-knowing” to being a source of essential skills, serving as a starting point for more advanced capabilities.

2. Pre-trained Data: From a “Hodgepodge” to Targeted Feeding

Previously, pre-training data consisted of a mix of various sources, with web pages accounting for 85% of the data. Now, the proportion of web pages has dropped to 15%, and code has become the primary source of training data. This change is due to model companies having a clear understanding of the specific capabilities required by their models (such as programming, reasoning, and handling complex tasks), so they provide targeted data accordingly.

Another important development is the increasing use of synthetic data, which is artificially designed to fit the needs of the training tasks. For example, a piece of code might be modified to include elements like bug finding, modification, tool invocation, and sequential task completion, allowing the model to learn about real-world scenarios during pre-training. This approach is more practical than simply memorizing words in the past.

3. RL and Post-training: From Fine-tuning to Skill Enhancement

Post-training used to be limited to making minor adjustments (e.g., improving the model’s chat capabilities), but now RL has become a crucial factor in enhancing model performance. If RL is done well, the pre-trained model can significantly improve its abilities. Some industry experts even suggest that the amount of computing resources invested in pre-training and post-training is now comparable.

This indicates that the ultimate capabilities of a model no longer rely solely on pre-training; post-training and RL are the key determinants of its performance. It’s like buying a rough house (the Base Model) and then spending a significant amount on renovations to make it habitable.

4. Mid-training: A Buffer Zone between Pre-training and Post-training

In the past, there was a stark contrast between the static data used in pre-training (web pages, books) and the dynamic tasks required in post-training (reasoning, agent interactions, tool usage). This mismatch often caused compatibility issues, especially with multi-expert models where the division of labor was fixed. Mid-training serves as a transitional phase, exposing the model to long contexts, reasoning data, and task scenarios before it moves on to more complex tasks. It’s like transitioning directly from elementary school to university; Mid-training acts as the middle school, helping the model make a smooth transition.

5. Simplified Training Paradigm: “Two Steps” to Train Large Models

The entire training process can now be summarized in two main steps:

  • Step 1: Supervised Learning (pre-training + Mid-training): The model learns foundational skills and knowledge, such as understanding words, Python, and logical concepts.
  • Step 2: Reinforcement Learning (RL): The model is exposed to real-world scenarios where it can learn by trial and error, combining its basic skills to solve complex problems (e.g., writing software, making decisions).

The boundaries between the various training phases are becoming increasingly blurred, with the core idea being to lay a solid foundation before applying those skills—just like learning basic arithmetic operations before moving on to solving more advanced problems.

Conclusion

The title “The Base Model Is Dead” is an exaggeration, but the role of the Base Model has indeed shifted from being the central focus to serving as a starting point for model development. It remains crucial, as without the necessary skills, RL cannot be effectively utilized. The industry’s shift focuses on how to apply basic skills through RL to create more powerful models. This represents a critical transition in the large-model field from emphasizing knowledge acquisition to focusing on practical application capabilities.