虎嗅

Gemini 3.5 Pro faces delays in release; Google has already submitted two preliminary drafts.

原文:Gemini 3.5 Pro难产,谷歌先交了两张草稿纸

Summary of Key Points

Google's flagship AI model, Gemini 3.5 Pro, which was originally scheduled to be launched in June, has been significantly delayed. Instead, in July, the company suddenly released two "lightweight" models focused on efficiency and cost: Gemini 3.6 Flash (emphasizing Agent capabilities and low cost) and Gemini 3.5 Flash-Lite (focusing on ultra-fast performance). Additionally, a security-oriented model, Gemini 3.5 Flash Cyber, was quietly launched, available only to specific customers. The delay in the flagship model has raised doubts about Google's AI capabilities, while the new models reveal a shift in strategy: from competing on the basis of having the smartest models to making AI more affordable and scalable, in an attempt to bridge the gap caused by the delay.

Detailed Analysis

1. Delay in the Flagship Model: Google's Anxiety About Its "Hard Power" Becomes Obvious

Gemini 3.5 Pro was intended to be Google's proof of its leading AI capabilities, with a launch planned for June. As of mid-July, there has been no news about it. According to Bloomberg, the release has been postponed for several months, and Google is urgently working on optimizing its coding abilities—a core competency in the Agent era, such as the ability of AI to modify real code and complete complex development tasks. Why are coding abilities so important? Because businesses now need AI to perform "real work," such as automatically writing code and operating computers to complete cross-application tasks, all of which rely on coding. If the flagship model fails to meet these standards, Google will struggle to compete with companies like OpenAI (GPT series) and Anthropic (Claude series).

The market's question is straightforward: With access to top research resources like DeepMind, why can't Google deliver a competent flagship model? Wall Street and the industry are waiting to see if Google can keep up with its competitors.

2. Gemini 3.6 Flash: Not Necessarily Smarter, but More Affordable

This model is designed for everyday AI tasks in businesses, such as customer service robots, automatic data analysis, and Agent assistants that require multiple AI calls. Its improvements are all focused on efficiency:

  • Lower cost: The output cost has been reduced from $9 per million tokens for Gemini 3.5 Flash to $7.5, saving 16% per call. Internal tests show a 55.8% reduction in token consumption for the same task (tokens are the units of text processed by AI; the fewer tokens used, the more cost-saving).
  • Faster performance: Task completion time has been halved from 2.7 minutes to 1.3 minutes.
  • Enhanced Agent capabilities: For example, scores in computer operation tasks have increased from 78.4% to 83%, and code testing scores have risen from 37% to 49%.

However, its overall intelligence has not improved (third-party test scores are the same as those of the previous generation). In simple terms, it may not be the smartest model, but it is more cost-effective for businesses that use AI extensively.

3. Gemini 3.5 Flash-Lite: Ultra-Fast Performance Without Price Reduction

This is the "fast and lightweight" version of the Flash series:

  • Unmatched speed: It can generate text at a rate of 350 tokens per second (equivalent to writing over 200 Chinese characters per second), nearly three times faster than the previous generation, making it one of the fastest models available.
  • Outperforms its predecessor: Although it is a lightweight version, some test scores surpass those of the more expensive Gemini 3.5 Flash. For example, its score in software engineering tests is 54.2% compared to 49.6%. This means that tasks previously requiring larger models can now be handled with this cheaper model.
  • Only downside: The price remains unchanged. Compared to the previous version (3.1 Flash-Lite), both input and output costs have not decreased. However, due to improved efficiency, it may indirectly save on resource usage.

This model is suitable for real-time chat robots, customer service that requires immediate responses, or automated tasks with high speed requirements.

4. Google's Strategic Shift: From "Showing Off" to "Being Practical"

Previously, Google competed by showcasing the intelligence of its models (e.g., handling large amounts of context or performing complex reasoning). Now, the focus has shifted to making AI more affordable and scalable. CEO Sundar Pichai has repeatedly emphasized that businesses can save $1 billion annually by using Gemini for their AI tasks. The reason for this shift is that what companies need is AI that is practical and accessible. For example, a customer service system that calls AI 1 million times a day could save $3.65 million annually if each call saves just 1 cent. By releasing these two Flash models, Google aims to capture the market for large-scale AI applications.

However, the delay in the flagship model raises concerns: Could the use of the Flash models be seen as Google avoiding more challenging tasks? If Gemini 3.5 Pro ultimately proves to be less capable than its competitors, even the lower-cost models will not be sufficient, as the flagship model represents true "hard power."

5. The AI Race Becoming More Competitive: Is Google's Lead at Stake?

The AI landscape used to be dominated by three global giants (OpenAI, Anthropic, and Google DeepMind), but domestic models like GLM-5.2 and Kimi K3 are also making significant progress. For instance, Kimi's long-context capabilities have caught up with Google's. Google's current situation is somewhat awkward: its flagship model is delayed, while competitors are advancing, and domestic models are catching up. If Gemini 3.5 Pro does not launch soon or performs poorly, the market may question whether Google is "big but not strong."

Next week, Google will release its second-quarter financial report, and investors will certainly ask about the progress of Gemini 3.5 Pro. This is not just about a single model but also a test of the reliability of Google's AI strategy.

In Conclusion

Google has temporarily stabilized its position with these two efficiency-focused models, but the delay in the flagship model exposes issues with its development pace. To maintain its leading role, Google must deliver tangible evidence of its AI capabilities as soon as possible. Otherwise, even the most affordable Flash models will not be enough to support its ambitious AI goals.