Summary of Key Findings
In the past half-month, a plethora of open-source large models have been released both domestically and internationally (domestically: Kimi K3, Qwen3.8-Max, DeepSeek V4 Pro; internationally: Meta Muse Glimmer, NVIDIA Nemotron3.5 Lightning). The author tested these five models in the same environment across five tasks: logical reasoning, code visualization, agent-based problem-solving, Blender 3D modeling, and writing. The results are as follows:
- The Overall Champion: Kimi K3 (scored full marks on six out of ten tasks).
- The Cost-Effectiveness Queen: DeepSeek (strong capabilities at an extremely low price).
- The Stable Performer: Qwen (can complete tasks, but with slower thinking processes).
- Limitations of Smaller Models: Muse and Nemotron perform well on simple tasks but struggle with more complex ones.
The open-source ecosystem is becoming increasingly vibrant, offering users a wider range of options.
1. Test Results: Who Is the Ultimate Champion, and Who Shines Due to Cost-Effectiveness?
In this test, Kimi K3 emerged as the top performer across all tasks—acquiring full marks in logical reasoning, accurately generating subway map web pages with transfer animations, and successfully creating Blender 3D scenes. However, it has significant drawbacks: it costs $0.35 per use (eight times more expensive than DeepSeek) and is also slower.
DeepSeek was the biggest surprise: Although its capabilities are not weak (it solved all password lock puzzles correctly without any limits), its price is incredibly low—just $0.40 for five additional questions in a time extension, and only $0.30 per bug fix. This means it can handle complex tasks at a fraction of the cost.
Qwen represents the “slow but steady” approach: it can solve many difficult problems given enough time (especially when output restrictions are removed), but its long thinking processes often result in interruptions, making it suitable for less time-sensitive scenarios.
Smaller models like Muse and Nemotron are more limited: they are fast and inexpensive for simple tasks (such as generating a 24-hour schedule), but they struggle with more complex ones. For example, they failed to solve password lock puzzles or correctly generate subway map routes, and encountered issues with bill processing due to coding limitations.
2. Strengths and Weaknesses of Each Model: No Model Is Perfect, But They All Have Their Advantages
- Kimi K3: A versatile model suitable for high-quality tasks. It excels at generating SVG images (e.g., recognizing a panda riding a bike with an umbrella), creating subway map web pages with detailed transfer animations, and building 3D scenes in one go. However, it is expensive and slow.
- DeepSeek V4 Pro: The cost-effective champion, ideal for scenarios where cost is a concern. It performs well on tasks like bug fixing and bill processing, and the five additional questions in the time extension only cost $0.40—four times less than trying with Muse. Its only downside is that Blender scripts occasionally use outdated parameters, causing errors.
- Qwen3.8-Max: A stable model suitable for patient users. It completes tasks well when given more time, but its long thinking processes can lead to interruptions.
- Muse Glimmer (Meta): A smaller model designed for simple tasks. It can handle the 24-hour schedule task, but it struggles with password lock puzzles and produces poorly readable SVG images. Bill processing also encounters coding issues.
- Nemotron3.5 Lightning (NVIDIA): Another smaller model suitable for basic tasks. It is fast for the 24-hour schedule task, but it failed to solve the password lock puzzles even after multiple attempts and lacks certain features in the subway map generation task. Its writing output also has an obvious AI presence.
3. Large Models vs. Smaller Models: Not Substitution, But Complementation
The test clearly shows that large and small models have different roles and do not replace each other:
- Large Models (Kimi K3, DeepSeek, Qwen) are ideal for complex tasks involving multiple steps of reasoning, handling complex code, sequential tool usage (bug fixing), and 3D modeling. Smaller models either cannot perform these tasks or do so poorly.
- Smaller Models (Muse, Nemotron) are better for simple tasks such as basic calculations and straightforward text modifications. They are fast and inexpensive, making them a waste of resources when used for more complex tasks.
For example, in the time extension, Nemotron spent less than $10 on ten attempts at the 24-hour schedule task, but failed to solve the password lock puzzles; DeepSeek, as a large model, solved all five questions correctly in one attempt, although it is more expensive.
4. The Impact of the Surplus of Open-Source Models
The release of so many open-source models in the past half-month brings numerous benefits:
- More Choices: Users no longer have to rely on closed-source models like ChatGPT; they now have options like Kimi K3 and DeepSeek that suit different needs (e.g., high quality or cost-effectiveness).
- Improved Cost-Effectiveness: Models like DeepSeek are lowering the cost of using AI, potentially leading to better models at lower prices in the future.
- Technological Progress: Open-source competition drives continuous improvement. Alibaba has released the Qwen3.8 weights, and Meta has announced the upcoming release of Muse Spark1.2, accelerating technological advancements.
In short, the more open-source models there are, the more benefits users have—more tools available at lower costs.
Conclusion
This test demonstrates that open-source large models can compete with closed-source ones and each has its unique strengths. For both individual users and businesses, there are now higher-quality, more cost-effective options available. As competition intensifies, models will become better and more affordable. This wave of open-source development is truly good news for users.