虎嗅

DeepSeek’s “eye-opening” capabilities revealed: It can recognize Elon Musk, but not Leung Man-fung.

原文:实测“开了眼”的DeepSeek:认得马斯克,认不出梁文锋

Summary of Key Points

DeepSeek has finally added multi-modal capabilities (the ability to process images) and released the experimental model Vision-Exp. The pricing remains very affordable (one cent allows users to view 8 images), but user reviews are mixed: the model fails to recognize its own founder, Liang Wenfeng (mistaking him for Wang Xing or Zhang Yiming), confuses real-person combinations with AI-generated images, and performs poorly on complex tasks such as replicating web pages. Although official benchmarks show it is close to top-tier models, there is a significant gap between the theoretical performance and the actual user experience. By releasing this “semi-finished” product, DeepSeek aims to gather user feedback to prepare for the official release.

Highlights:

  • Image Processing for the First Time, and at an Incredibly Low Price

DeepSeek’s previous models could only process text. Vision-Exp now includes visual capabilities, such as identifying people, analyzing screenshots, and interpreting charts. The pricing is also attractive: each image costs at most 384 Tokens (the model’s unit of computation), allowing users to view 8 images for just one cent, even during peak usage times. This continues DeepSeek’s tradition of offering affordable solutions, making it accessible to both developers and ordinary users.

Issues Encountered by Users:

  • Misidentifications and Confusions

Users have encountered numerous amusing mistakes, such as:

  • Misrecognizing Founders: The model mistakenly identified DeepSeek’s founder, Liang Wenfeng, as Wang Xing from Meituan or Zhang Yiming from ByteDance.
  • Doubting the Authenticity of Real People: Users doubted the authenticity of photos of the “Times Youth Group,” claiming the members looked too similar and had too smooth skin, suggesting they might be AI-generated.
  • Poor Performance on Complex Tasks: When asked to generate a Three.js 3D scene, the model’s output was slightly better than Gemini 3.7 Flash but required significantly more Tokens and took much longer (15 times as much effort and 3.5 times as long).

These issues have led users to joking that the model is “blind” and even describe it as a “good model that’s still in its early stages.”

The Gap Between Benchmarks and Real-World Performance

Although official data suggests Vision-Exp’s multi-modal capabilities are on par with Anthropic’s top model, Opus, the actual experience is far from ideal:

  • Difference Between Lab and Real World: Benchmarks are conducted under standardized conditions, while users face images of varying quality (blurry, synthetic) and complex tasks (memes, complex web pages) that the model struggles with.
  • Insufficient Detail Handling: For example, when replicating web pages, Vision-Exp only provides a basic framework; colors and text positions are incorrect, whereas GLM-5.3 and Kimi K3 can capture more details.

Why a Semi-Finished Product?

The model’s name (“Exp” – Experimental) indicates that DeepSeek is aware of its limitations. Why release it so early?

  • Cost Control: Using the V4-Flash model (with fewer parameters and faster performance) as a foundation, DeepSeek can maintain an affordable price, encouraging widespread testing by developers.
  • Gathering Real-World Feedback: Issues that cannot be identified in lab tests (like misidentifying founders or misunderstanding memes) can only be uncovered through user usage, providing valuable insights for improvements.
  • No Interruption for Existing Users: Users can continue using the V4-Flash version, while those interested in trying the new model can do so without interference.

This is not the first time DeepSeek has adopted this approach: in September 2025, it released V3.2-Exp for testing new features, with the official version following two months later.

Future Prospects:

Vision-Exp’s current shortcomings highlight areas that need improvement, such as inaccurate face recognition and a lack of understanding of cultural references. DeepSeek needs to address these issues to make the official version more reliable and effective. Whether Liang Wenfeng can be correctly identified in the future version remains to be seen… Only time will tell how well these “eyes” (the model’s capabilities) will develop.

In summary, DeepSeek’s multi-modal model has potential, but it still needs significant refinement. The experimental version serves as a valuable learning opportunity; the success of the official release will depend on how well these issues are addressed.