虎嗅

"Kimi K3's First Test: A Close Battle with Opus4.8?"

原文:跟Opus4.8打的有来有回?,Kimi K3首测

Summary of Key Points

This article is the first test of the domestic large-scale model Kimi K3. The author compared it with the internationally renowned model Claude Opus 4.8 using four challenging scenarios from the front-end development field: 3D glass replication, WebGL light bands, parallax scrolling animations, and furniture web page effects. The results show that Kimi K3 can hold its own against Opus 4.8—K3 won one scenario (3D glass video replication), tied in two (Bezier light bands), and lost slightly in the other two (parallax scrolling and furniture effects). Kimi K3’s standout features include its multimodal capabilities (it can understand videos and generate 3D models) and its ability to self-correct errors, while Opus excels in layout aesthetics and the success rate of its initial outputs. More importantly, Kimi K3 is significantly cheaper than GPT-5.6-Sol and Opus 5, representing a significant advancement for domestic AI models.

Detailed Analysis

1. The Four Challenge Scenarios: A Close Battle Between K3 and Opus

The author selected scenarios that require high technical expertise in front-end development, not just simple web page creation:

  • 3D Glass Prism Video Replication: Kimi K3 won.

The model was tasked with replicating a rotating 3D glass image and replacing the text. Kimi K3 produced a clear, well-structured image with neatly arranged elements; Opus’s version had a blurry effect, with overlapping text. Kimi K3’s success lies in its ability to understand the 3D structure of the video.

  • Bezier WebGL Light Bands: The task involved creating glowing curves on a black background with glass-like dispersion effects and mouse interaction. Kimi K3’s version was clean and detailed, with neat light bands and additional visual effects; Opus’s version had more noticeable dispersion but coarser light bands and a brighter overall appearance. Both versions had their merits, so it was a tie.
  • Parallax Scrolling Cards: Nine cards needed to scroll through five different layouts (fan shape, grid, etc.) smoothly. Opus’s first screen displayed the title clearly with properly arranged cards; Kimi K3’s first screen had the title disappear suddenly, leaving only one card visible. Although the scrolling animation worked, the overall impression was less favorable.
  • Furniture Web Page Effects: The goal was to create a webpage that expanded upon hover. Both models achieved the desired interaction, but Opus used a dark theme that resembled a professional design studio website; Kimi K3’s bright style lacked depth and professionalism. Opus won in terms of aesthetics.

2. Kimi K3’s Unique Strength: The Ability to Understand Videos and Improve Code Automatically

Kimi K3’s most impressive feature is its multimodal capability. In one scenario, it was able to directly extract 3D shapes and glass materials from the video and generate code without any additional text instructions, something Opus could not do.

Another example is the light band generation: Kimi K3 initially produced a thick, stick-like version of the light band. After checking the result, it automatically corrected it to make it thinner. This self-correction ability indicates that the model can not only generate content but also evaluate its own work.

3. Opus’ Strengths: Consistency and Professional Aesthetics

Opus’s strengths include a high initial success rate and excellent layout aesthetics. For instance, in the parallax scrolling scenario, it produced the correct first screen from the first attempt; in the furniture effects scenario, its dark design, numbering, and vertical text made the webpage look more professional, like a real-world website. Opus’s approach ensures more stable results, making it suitable for scenarios where perfection is required from the first try.

4. Why Kimi K3 Is Worth Celebrating?

Despite being more expensive than its predecessor, Kimi K2.6, its parameters have significantly increased, and it remains much cheaper than GPT-5.6-Sol and Opus 5. For ordinary developers and businesses, this means they can access nearly top-tier AI models without incurring high costs. This is a significant advancement for domestic AI models, as not everyone can afford international alternatives.

5. The Technical Difficulty of the Tests

The scenarios chosen by the author are quite challenging even for human developers: for example, creating WebGL light bands requires writing GPU-specific code for pixel-by-pixel calculations, and implementing parallax scrolling involves coordinating the five different layouts of nine cards. Human developers might spend a whole day on such tasks. The fact that these models can generate the correct results with minimal adjustment shows how much time they can save.

Conclusion

Although Kimi K3 has not surpassed Opus 4.8 in all aspects, it stands out for its multimodal capabilities and self-correction abilities, as well as its affordable price. For users, this means they can obtain nearly world-class AI development assistance at a lower cost, which is a positive sign for the development of domestic AI models.