Summary of Key Points
At the 2026 World Artificial Intelligence Conference (WAIC), two significant trends emerged in the field of large-scale AI models:
1. Leading companies are collectively striving to achieve model sizes with trillions of parameters, a path that has proven feasible and capable of breaking industry monopolies.
2. There is disagreement regarding the multi-modal technology approach, but there is consensus that the real bottleneck for future AI lies not in the choice of modalities, but rather in breakthroughs in underlying capabilities such as learning paradigms and autonomous evolution.
Detailed Analysis
1. The Race to Trillions of Model Parameters
What drives the trend towards larger model sizes? Industry experts have identified three main reasons:
- The Scaling Law: There is a general belief that larger parameter sizes lead to better model performance, as repeatedly demonstrated by Anthropic's Claude series (from Sonnet to Opus 4 to Fable 5, with each increase in parameters resulting in significant improvements).
- Need for Agent Automation: AI systems are needed to autonomously complete various tasks similar to humans (e.g., writing reports, handling complex processes), and larger parameter models can better support this type of continuous work.
- Long-Term Task Requirements: AI has evolved from being capable of performing simple tasks for a few seconds to working independently for extended periods (as mentioned by Jiēduò Xīngqí). Larger parameters are essential for achieving this level of capability.
Example: After the release of the Kimi K3 model with 2.8 trillion parameters, MiniMax announced that its next-generation model, M3 Pro, will have over 2 trillion parameters. iFlytek also plans to develop a trillion-parameter model. Even companies that have not officially released new models consider achieving this scale as a standard for their next releases.
2. The Impact of Large-Parameter Models
Large-parameter models are more than just about increasing size; they can also address industry challenges:
- Breaking Monopolies: Gavin Baker, CIO at American investment firm Atreides, believes that the Kimi K3 may mark a turning point in the industry, as it allows more companies to participate in the AI market, which is detrimental to giants like Anthropic and OpenAI but beneficial for most enterprises.
- Clear Division of Labor: The industry has seen a shift towards models with different sizes: those with hundreds of billions of parameters are used for cost-effective, general tasks (e.g., chatting, simple document processing), while trillion-parameter models serve as “flagships” for tackling complex reasoning and replacing human-intensive tasks.
3. The Debate over Multi-Modal Approaches
There was significant discussion about multi-modal technologies (combining text, images, video, and sound) at the conference, with two main viewpoints:
- Language as the Core: Professor Qiu Xipeng from Fudan University argues that language is fundamental for advanced human reasoning and that multi-modal technologies complement rather than replace language, helping models understand real-world scenarios. He pointed out that the current issue with AI models is their lack of understanding of real contexts, suggesting a future combination of basic large models with multi-modal interaction tools.
- Multi-Modal as the Foundation: Zhang Xiangyu from Jiēduò Xīngqí emphasizes that human intelligence relies on multiple modalities (sight, hearing, touch), and single-language models are limited in their ability to replicate physical relationships (e.g., understanding that a glass will break when dropped). Liu Ziwei from Nanyang Technological University in Singapore compares language data to “fossil fuels” (quickly depleting) and multi-modal data to “new energy” for sustained AI development.
Xu Li from SenseTime believes that having multiple approaches allows for diverse innovation in the industry.
4. The Next Major Hurdle for AI
The real challenge is not in modalities but in learning paradigms and autonomous evolution:
Zhao Deli from Alibaba’s DAMO Academy highlights that current language models’ success is based on static, high-quality text data and mimicking human learning methods. In the era of dynamic intelligent agents, both language and vision models face the same problem—lacking the ability to learn and evolve from sparse environmental feedback (e.g., repeating mistakes).
The industry consensus is that the current “language-based + multi-modal” architecture has inherent flaws, and future breakthroughs will come in learning paradigms, memory mechanisms, and autonomous evolution capabilities.
In Summary
While AI models are working on increasing parameter sizes to break monopolies and tackle complex tasks, as well as exploring multi-modal approaches, the ultimate goal is to enable AI to learn and evolve autonomously, just like humans. The analysis uses clear language that makes the industry trends and key issues accessible to non-financial professionals.