Summary of Key Points
Three significant AI developments occurred in Silicon Valley overnight yesterday: Google released Gemini 3.8 Flash, Meta updated Muse Spark1.3, and the startup Mostik made its debut on WIRED. On the surface, these updates refresh the rankings of AI model capabilities, but they actually point to a “multi-dimensional revolution” in reducing the cost of AI inference. This revolution encompasses several aspects: from “making high-end capabilities more affordable” to “minimizing unnecessary tasks,” and even to “skipping natural language and directly transmitting internal data.” The economics of AI are shifting from being priced based on the number of tokens used to being priced based on the actual cost of the tasks being performed. As a result, intelligence is becoming increasingly accessible, which could potentially lead to the emergence of a large number of AI applications that were previously not viable due to high costs.
1. Google: Bringing High-End AI Capabilities to More Affordable Packages
Google’s Gemini 3.8 Flash, which was previously known for being fast and inexpensive but with limited capabilities, has now incorporated top-tier AI software engineering skills (such as completing complex code tasks in DeepSWE tests) into its lower-priced range.
- Performance: It achieved 74% on the DeepSWE test, on par with the most advanced model, Claude Opus5.
- Cost: Completing a task costs an average of only $2.36, compared to $11.84 for Claude and nearly $6.46 for GPT.
Why is this important? In the era of AI agents, tasks are not just about simple conversations; they often involve multiple steps (such as writing code or conducting research). The significant cost difference can influence whether companies are willing to adopt these technologies. Google argues that with lower token prices, models can “think more” (i.e., perform longer inference processes) without having to compromise on efficiency to save costs.
2. Meta: Fewer Mistakes in AI = Direct Cost Savings
Meta’s Muse Spark1.3 has seen its performance improve from 55% to 75.4%, but the more crucial figures are the 20% reduction in the number of tool calls and the 25% reduction in token consumption.
- Cost Savings: The biggest waste in AI agents is making mistakes—for example, an AI that misunderstands requirements and rewrites multiple files or runs unnecessary tests, wasting thousands of tokens. Muse can now identify ambiguities earlier, request help when needed, and remember initial constraints, thus avoiding unnecessary efforts.
- Implication: In the era of chatbots, improving scores by 20% might seem impressive, but in the era of AI agents, making fewer mistakes is more directly beneficial in terms of cost savings.
3. Mostik: AI Models Communicate Without Natural Language
Mostik has taken a counterintuitive approach by allowing two AI models to communicate directly using “mathematical signals” rather than natural language.
- Comparison: Currently, models communicate by converting data from one format to another (e.g., printing a file and scanning it). Mostik eliminates this intermediate step, allowing the models to transfer data directly.
- Experimental Results: By bridging a large model (GLM-5.2 with 753B parameters) with a smaller model (Qwen3.5 with 4B parameters), Mostik achieves performance between the two while reducing costs by 20%.
- Challenge: Different models use different “languages” (architectures and parameters), making communication difficult. However, a 20% reduction in costs is already very promising.
4. The Future: Tiny Models on Mobile Devices, Complex Tasks on the Cloud
In the past, efforts were focused on reducing the size of large models for mobile devices (from 7B to 4B to 1B parameters). If Mostik’s approach proves successful, future phones might only need tiny models with a few tens of megabytes of parameters.
- Task Distribution: Local models will handle simple, frequent tasks (e.g., understanding user activity and current state), and complex tasks will be processed on the cloud. The cloud will then send the necessary signals back to the mobile device, eliminating the need to transfer large amounts of data (in the form of tokens).
- Impact: This architecture could revolutionize AI computing, reducing costs by an order of magnitude. What currently costs dozens of dollars per task could potentially cost just a few cents.
Conclusion
The real breakthrough will occur when intelligence becomes so affordable that it can be used freely. The common theme of these developments is that the cost of AI is not decreasing linearly but is being fundamentally restructured. When intelligence is cheap enough to perform tasks effortlessly or to have multiple AI agents working around the clock without incurring significant costs, many currently impractical applications (such as personalized AI teams or fully intelligent assistants) will become commonplace. Rather than focusing on which model ranks first, it’s more important to consider how much cheaper these intelligent systems can become in the long run.
(The entire analysis is written in plain language to make the logic and impact of the AI cost revolution understandable to non-financial and non-AI professionals.)