Summary of Key Points
This article challenges simplified notions such as "LLMs represent high-end manufacturing" or "AI's ultimate goal is power generation," emphasizing that the essence of the LLM competition lies in the competition of mathematical和应用 mathematical skills alongside engineering intuition. Chinese teams like DeepSeek and Kimi have broken the outdated narrative that "China can only excel in engineering but not in innovation" by developing original technologies (such as MLA compression and KDA long-range reasoning), thereby reshaping the global asset pricing logic. It also highlights that LLMs currently lack a solid mathematical foundation and face many unresolved theoretical issues. In the second half of the competition, greater emphasis should be placed on mathematical talent and intuition; otherwise, there is a risk of making unguided, reckless progress.
1. Why LLMs Are Not "High-End Manufacturing?"
The core of manufacturing is the combination of resources and process optimization—for example, Elon Musk has the necessary personnel, chips, and supply chain, but even his Grok model performs mediocrely, and he still has to rent tens of thousands of chips from Anthropic. LLMs, on the other hand, cannot be improved simply by buying more hardware or reorganizing structures; the key lies in a deep understanding of the underlying mathematical principles. For instance, the DeepSeek team discovered that the KV Cache (which stores critical information when processing context) appears to be high-dimensional and complex, but it actually has a low-dimensional structure—similar to a collection of points that all fall on the same plane. This kind of insight is not the result of engineers trying to find compression methods; it is a product of the intuition in physical mathematics to identify simplifications in complex systems, something that cannot be replicated by manufacturing processes.
2. How Chinese Teams' Original Breakthroughs Have Changed the Narrative
The market previously believed that Chinese AI could only perform engineering tasks without original innovation, but DeepSeek and Kimi's technologies have changed this perception:
- DeepSeek's MLA technology: This uses mathematical methods to decompose the KV Cache into a lower-rank structure, significantly reducing model costs, and the framework has been made open-source. Without this, domestic models like Zhipu might still be limited to public open-source frameworks.
- Kimi's KDA technology: This redefines the attention mechanism, solving the problem of error accumulation in long-text reasoning (e.g., maintaining context over multiple pages).
These breakthroughs have not only helped domestic LLMs catch up with the leading players but also challenged the monopoly of OpenAI and Anthropic. They have also changed the narrative about Chinese AI, as evidenced by reports on the official website of the Chinese Academy of Sciences.
3. The Importance of "Intuition" for LLMs
Here, "intuition" refers to the ability to see the essence in complex systems. For example, the KV Cache in standard attention mechanisms is a high-dimensional matrix, and while graduate students familiar with linear algebra know about low-rank decomposition, only the team led by Liang Wenfeng realized that the KV Cache has a low-dimensional structure in the sequence dimension. Long-range reasoning involves more than just adding additional components; it requires a reconstruction of the mathematical framework to prevent error accumulation and maintain semantic coherence. This kind of intuition is not learned through training but is a natural ability. Examples include Kenji Ganley, the father of information geometry, who independently proposed the natural gradient descent method, and Hideo Fukushima, who invented the precursor to CNN. The success of these individuals in China is a result of their talent and the right environment.
4. The Mathematical Shortcomings of LLMs and Their Hidden Risks
Current LLMs are similar to the communications industry before Shannon: engineers could build telephones, but no one understood the essence of information. Many key issues in LLMs lack rigorous mathematical proof:
- Why does the Scaling Law (better models with more parameters) follow a double-power law? Will it still hold true with trillions of parameters? We don't know; we'll have to wait and see.
- Why does the GRPO algorithm (used by DeepSeek) converge? The team only says it hasn't failed, but later studies revealed that it has unavoidable biases.
- The "illusion problem" occurs when models generate random outputs, which is due to the lack of a mathematical framework to guide them to make optimal decisions between correct answers and admitting ignorance.
If these issues are not resolved (e.g., if the Scaling Law stops working with larger parameters, or if compression and prediction techniques are merely superficial), all valuations based on the idea of AGI (Artificial General Intelligence) could collapse, leading to market turmoil.
5. The Future of the LLM Competition: Focusing on Mathematics
The article suggests that mathematics will become even more crucial in the future of LLM competitions. Current LLMs lack fundamental mathematical foundations. To overcome these challenges, we need to:
- Recruit mathematical experts: People skilled in non-convex optimization, matrix calculations, and differential equations to develop the necessary formulas.
- Wait for the "Shannon moment": Just as Shannon defined information with a few formulas, someone in the future will need to provide concise formulas to understand the limits of LLM performance, the lower bounds of reasoning consistency, and whether self-regressive errors can be mitigated.
Only then can LLMs evolve from random generators to truly reliable tools. Otherwise, the high-profit narratives around MaaS (Model as a Service) will be unfounded, and no one will be willing to pay for unproven technologies.
Final Advice for Young People
The author encourages talented and intuitive individuals to study mathematics and physics. Perhaps one day, someone will develop the "Shannon theorem" for LLMs, defining the boundaries of equivalent intelligence with mathematical formulas. That would be a true breakthrough that money cannot buy.
The core message of this article is that LLMs are not about accumulating resources but about discovering the underlying principles. Mathematics and intuition are the keys to success, and Chinese teams have taken a significant step forward, but there is still a long way to go.