Summary of Key Points
Google has released three new Gemini Flash series large models in just six weeks. The latest model, 3.8 Flash, remains incredibly affordable, with an input cost of $0.75 per million tokens and an output cost of $3.75. Its performance scores are impressive: it has matched Claude Opus 5 in programming capabilities and surpassed GPT-5.6 Sol in reasoning tasks. Additionally, a Cyber version specifically designed for cybersecurity has been introduced. However, in real-world usage, ordinary users have found that while the model scores well on tests, its practical application is subpar. This is due to the fact that Google’s flagship model, Gemini Pro, is facing delays in development, and there has been a loss of core technical talent. As a result, Google is relying on the Flash series to maintain its presence in the market by making rapid, small-scale advancements.
1. The Affordable “Performance Star”: Performance Close to Flagship Level
3.8 Flash is described by Google as the “smartest” Flash model, with a focus on enhancing long-term programming, complex reasoning, and professional agent scenarios, yet its price remains the same as its predecessor, 3.7 Flash.
- Programming Abilities Approaching the Limit: In the DeepSWE long-term programming test, it scored 73.7%, matching Claude Opus 5 (74%), but at a fraction of the cost (2.36 dollars per task compared to 11.84 dollars for Opus 5). In the Terminal-Bench 2.1 test, it scored 89.4%, bringing Flash models closer to the 90% mark for the first time.
- Interdisciplinary Reasoning in the Top Tier: In the HLE (Human Legacy Exam) that covers STEM and humanities, it scored 54.9%, slightly higher than Opus 5 (54.4%) and GPT-5.6 Sol (54.5%).
- Professional Scenarios Surpassing the Flagship: In the financial agent test (Vals V2), it scored 61.4%, exceeding GPT-5.6 Sol (53.8%), and in the legal agent test (Harvey), its pass rate was 10%, which is 1.5 times that of Opus 5.
- Cyber Version for Security: It has a vulnerability detection rate of 86.2% (higher than GPT-5.5-Cyber), and the cost of fixing vulnerabilities is lower. The Chrome team used it to generate correct patches, which were 2.6 times more effective than those produced by other models, but these patches are only available to “trusted defenders”.
The secret to its performance lies in its ability to think more strategically when faced with challenging tasks. For example, in programming tasks, 3.8 Flash generates 36,000 more tokens and executes 41 more steps than 3.7 Flash. Although the actual cost for complex tasks has increased slightly (from $0.4 to $0.58), it is still much lower than that of the flagship models.
2. Weaknesses in Real-World Tests: Good Scores, but Poor Practicality
While the official demonstrations show 3.8 Flash capable of building 3D structures and using a DOS version of Google Maps, users encounter significant shortcomings in practical use:
- Double Standards in Configuration: Official demonstrations use the Antigravity platform and Nano Banana to generate textures, while ordinary users only have access to basic models. For instance, building an Airbus helicopter with Three.js takes 109 seconds, but the resulting model has rough components and awkward movements; the 3D volcano simulation also has visual flaws.
- Fast but of Poor Quality: In the Sticky Ball game, Flash performs quickly, but its movement mechanics and gameplay are inferior to those of Kimi K3. In the New York city market scenario test, only 6 trees are included (compared to 10 in 3.7 Flash), and important details such as crosswalks and underwater effects are missing. Additionally, there are unnecessary camera presets and post-processing settings.
- Lack of Endurance for Long Tasks: In the Terminal-Bench 4.0 test (a more difficult task), it scored only 19.1% (compared to Opus 5’s 51.8%); in the OSWorld computer operation test, it scored 59% (compared to Opus 5’s 75.4%). Overall, it ranks eighth in the intelligence leaderboard, with strong performance in certain code-related tasks, but its capabilities fail in multi-domain, long-term tasks.
3. The Urgency Behind the Rapid Updates: Delays with the Flagship Model and Talent Loss
Google’s haste in releasing the Flash series is not due to a technological breakthrough, but rather due to pressing circumstances:
- Delays with the Flagship Model: In May of this year, Sundar Pichai announced that the new Pro model would be available in a month, but the 3.5 Pro plan was later canceled, as its improvements were not as significant as those of Flash. Gemini 4 is still in the later training stages and is far from completion.
- Loss of Core Talent: In the past one and a half months, key figures such as Noam Shazeer (a pioneer in large models) and Jeff Dean (a central figure in Google’s AI efforts) have left the company, causing organizational instability. The new management team is pushing for faster releases.
- Flash as a Temporary Solution: The Flash series is smaller in scale, and its modification and training requirements are much lower than those of the Pro series. Multiple internal teams can work on different versions simultaneously, allowing for rapid iteration every three weeks. In contrast, making adjustments to the Pro series requires more GPUs and more time, resulting in higher costs for experimentation.
Therefore, the release of three Flash models in six weeks is a strategic move to maintain Google’s competitiveness in the face of challenges such as delays with the flagship model and talent loss. It’s a way to keep up with competitors like OpenAI and Anthropic.
4. The True Role of Flash: Not the Ultimate Goal, but a Temporary Measure
Flash is not Google’s ultimate goal; it is a temporary solution. The issues identified in real-world tests indicate that it is not yet a fully capable flagship model, with weaknesses in long-term tasks and detail handling. However, it is affordable, can be updated quickly, and performs well in specific scenarios (such as programming and financial agents). This allows Google to maintain its market share and attract developers until the Pro and Gemini 4 models are ready.
In summary, the Flash series represents a “cost-effective alternative” for Google: it performs well in certain areas (programming, finance) at a lower cost, although it falls short in more comprehensive tasks. The rapid updates reflect Google’s struggle to develop a flagship model while dealing with internal issues and talent shortages. For users, Flash is a viable choice for specific use cases, but for more complex tasks, they will need to wait for the more advanced models. For the industry, this highlights the intense competition in the large model space, where no company can afford to slow down, even if it means making small, rapid advancements.