虎嗅

After Kimi K3, is Chinese AI entering the “Gaokao era”?

原文:Kimi K3之后,中国AI开始进入”高考时代”?

Summary of the Key Points

The domestically developed large model, Kimi K3, has ranked first in global programming rankings, but its performance in real-world development scenarios (such as writing functional BAT game scripts) is poor: the free version cannot control characters using a keyboard, and even the paid version fails to display any graphics; in contrast, GPT-5.6 Codex completed the task effortlessly. This highlights an issue in the AI industry—the excessive focus on “scoreboarding” (testing performance), while neglecting practical work capabilities. Behind this is the competitive pressure driven by capital, reminiscent of the subsidy wars in the internet sector. The article calls on AI companies to break out of this ranking-driven cycle and focus on meeting users' real needs and solving practical problems.

1. Why Can’t the World’s Number One Kimi Create a playable game?

Kimi K3 scored 1679 points, ranking first in the WebDev Arena, which sounds impressive. However, tests by friend Tong Tong revealed its shortcomings: when asked to create an ASCII version of “Bee Bee” using BAT scripts that required keyboard control and enemy movement, the free version of Kimi could display the game graphics but failed to handle keyboard input. Even after upgrading to the 699 yuan subscription version, it still couldn’t display the graphics correctly after five attempts. GPT-5.6 Codex, on the other hand, generated a functional program in one go.

This doesn’t mean that Kimi K3 is entirely ineffective; rather, it shows its limitations in tasks that require specialized training data and practical problem-solving skills, rather than just memorizing solutions.

2. Has the AI industry contracted the “college entrance exam syndrome”?

The article compares the current state of AI development to China’s education system: model companies are aggressively competing on various rankings (such as WebDev Arena and SWE-Bench), much like students preparing for the college entrance exam, with scores continuously rising, exciting both the media and investors. However, in real-world work, these models often fail—code doesn’t run properly, solutions look good but are difficult to implement, and beautiful PPTs contain many flaws.

Isn’t this similar to a student who scores 700 on the college entrance exam but struggles with tasks like creating PPTs, communicating with clients, or leading a team? The reason is that exams (rankings) and real work (solving practical problems) are fundamentally different. Rankings have standard answers, but the real world does not.

3. Why do companies obsess over rankings?

The reason companies focus on rankings is that they have become more of a reflection of financing performance than technical capability. In the past six months, China’s AI industry has entered a “monthly update era,” with models like Kimi, Tencent’s Hunyuan, and Alibaba’s Qianwen being released monthly. The competition is not about technology but about capital: higher rankings lead to higher valuations (e.g., DeepSeek’s valuation of $7.1 billion), making it easier to raise funds, which in turn allows companies to purchase more GPUs, hire more staff, train new models, and achieve even higher scores.

This is similar to the subsidy wars between Didi and KuaiDai or Meituan and Mobike a decade ago, except that the focus has shifted from subsidies to rankings. Companies are creating stories for investors rather than solving real problems for users.

4. Pioneers breaking the pattern

Not all companies are obsessed with rankings. For example, when Tencent released Hunyuan Hy3, it didn’t emphasize rankings but focused on practical use cases. OpenAI has also shifted its focus from rankings to more practical projects like Codex (which can complete entire programming tasks) and Agents (that can handle tasks like humans).

These approaches may not seem groundbreaking, but they are closer to real-world work. Real programmers spend only 20% of their time writing code; the remaining 80% is spent debugging, fixing bugs, and adjusting solutions. The essence of AI is not about writing code for the first time but about completing tasks effectively—just like Codex, which created a playable game without focusing on the beauty of the code.

5. The real battlefield for Chinese AI: user needs

The article acknowledges Kimi K3 as a significant advancement for domestic models (being the first to enter the global top tier) but warns against repeating the same patterns of competition seen in the internet sector. The focus has shifted from traffic, subsidies, and financing to parameters, rankings, and valuations, but the underlying drive remains the same.

True success lies in whether users are willing to use the technology daily. WeChat and TikTok succeeded not because of their rankings but because they met users’ real needs. The value of AI is not in taking exams on behalf of humans but in solving problems for them—e.g., writing and debugging code, or creating practical solutions that can be implemented.

In conclusion

Kimi K3 represents a milestone for domestic AI models, but winning rankings does not guarantee success in the long run. AI companies need to break free from the capital-driven ranking cycle and focus on solving real problems for users. This is where true competitiveness lies.