Summary of Key Points
This is a test report from a tech blogger on the newly launched V4.1 Flash model by DeepSeek, which is currently in beta testing. The new model features a completely new architecture and native multi-modal capabilities. The developers even directly asked users in the beta survey whether it could replace the previously more advanced V4 Pro. The blogger compared the new model with the old V4 Flash by running 14 sets of tasks, consuming a total of approximately 300 million characters in model calls. The results showed that the new model has significantly improved abilities in image recognition and visualization. However, it also has some drawbacks, such as fast typing but slow task completion, difficulty in remembering long pieces of content, and increased costs for more complex tasks, indicating that it is not a “perfect upgrade” that outperforms the old version in every aspect.
---
Easy-to-Understand Explanation by Dimension
1. The Most Prominent Improvement: It Finally Has “Its Own Eyes”
The old version of Flash was essentially a pure-text-based model. When you provided it with an image, it couldn’t understand it on its own and had to rely on third-party text scanning and image recognition tools, which could take dozens of seconds to minutes, and often resulted in incorrect interpretations (for example, it might misidentify a scalper on a poster or mistake English words in the image for unrelated ones). The V4.1 Flash model, being native multi-modal, can understand the content of an image in just 5-7 seconds, eliminating the need for external tools. It can even assess the quality of the image it creates (for instance, it could recognize mistakes in a drawing, such as a pelican riding a bicycle being depicted as a swan with a bowl). Additionally, it can directly transform the rules from the image into executable code. For example, given three reference images, the new model could create a game where an otter pushes a box, accurately reproducing the blue otter’s clothes, orange backpack, and the layout of the game board, without any additional steps.
2. Three Times Faster on Paper, but Why Isn’t the Final Result Much Faster?
Although the V4.1 Flash can type more than 200 characters per second, which is three times faster than the old version, the overall task completion time isn’t much faster. This is because the model spends a lot of time “self-checking” and revising its code. For example, after writing code, it runs tests to find and fix bugs, which takes significantly more time than the old version. It’s like hiring an intern who types three times faster than a regular person to create a PowerPoint: after finishing, the intern has to repeatedly review the content and ask three colleagues for help, resulting in a final delivery time similar to that of a slower typist. The speed advantage on paper is thus diminished due to these additional checks.
3. Severe Skill Imbalance
The new model’s strengths and weaknesses are very clear. It excels at turning ideas into actionable content (e.g., creating visual representations of twin prime numbers or 3D armed aircraft), but its memory for long pieces of text is extremely poor. For instance, when asked to write a 100,000-word article, it might mix up details from different sections. When creating a stock market report, it might show incorrect figures. This means it’s great for quick, visual tasks but not for writing long, detailed documents.
4. The New Model Is More Expensive?
Surprisingly, the new model is 36% more expensive than the old version for the same set of tasks (62 yuan vs. 45 yuan). The blogger found out that the extra cost comes from the model’s use of additional AI agents to handle complex tasks. For example, it assigns dedicated writers for each chapter of a 100,000-word article and uses multiple AI agents to write different parts of a large survival game. While this may lead to better-quality results, the increased cost is due to the additional resources used. Fortunately, the developers have announced a price cut on September 10th, reducing the cost for frequently used tasks by 60%.
---
Please note that this translation is adapted to fit the target audience and cultural context, maintaining the original structure and tone of the financial news analysis.