Summary of Key Points
Domestic large-model manufacturer DeepSeek, which had recently increased the prices of its API interfaces by up to 11 times in August, sparking widespread complaints and causing many developers to switch to other services, announced less than a month later that it would lower the prices of its flash series of models, which are known for their cost-effectiveness. Some of the fee items have been reduced back to the levels before the price increase. This price adjustment is not a result of DeepSeek’s “conscience awakening” nor a complete return to the old prices; rather, it is a strategic move in light of the upcoming launch of a new generation of low-cost models. The goal is to maintain high profits on its premium products while regaining the market share among lower-budget developers. This also marks a significant shift in the pricing competition for domestic large models, as they are moving away from the reckless spending phase and towards a more pragmatic approach to generating revenue.
---
Detailed Analysis
1. Understanding the Price Adjustment
The price reduction is not a complete return to the pre-increase levels but a targeted adjustment:
Many people might think that DeepSeek has caved in and refunded all the extra fees. However, this is not the case. Let’s put the numbers into context that would be easier for the general public to understand: 1 million tokens is roughly equivalent to the amount of information contained in a medium-length novel.
- Only two fee items have been reduced back to the pre-increase levels:
- “Cache hit input” (when the AI can answer a question based on previous knowledge, avoiding re-computation) has been reduced from 5 cents per million characters after the price increase to 2 cents, a 60% decrease.
- “Cache miss input” has been reduced from 1.5 yuan to 1 yuan, a 33% decrease.
- The fee for generating content, which was 2 yuan per million characters before the price increase and rose to 4 yuan in August, has not been changed and remains at 4 yuan.
Overall, the usage cost for average developers should decrease by about 40%, but the actual reduction will depend on the specific use case. If similar questions are frequently asked, the savings will be greater; if each request is unique, the reduction will be smaller. Additionally, during peak hours, the cost doubles, meaning that using the service during working hours will result in a higher overall cost.
2. The Price Reduction Is Not a Result of Goodwill but a Response to a Mistake
The August price increase was implemented alongside the launch of DeepSeek’s highest-end V4 Pro model, which saw prices increase by up to 11 times. This move significantly impacted small AI startups and individual developers that relied on the low-cost models, leading many to switch to alternatives from ByteDance or OpenAI. DeepSeek could not afford to lower the prices of the V4 Pro model, as it is designed for large companies with substantial budgets seeking top-tier performance. Therefore, the price reduction was limited to the more affordable flash series, offering a compromise to those who had left: “We’re keeping the cheaper options as affordable as before, while maintaining our high-end products’ profitability.”
3. Cost Reductions Are Possible Because New Models Have Lowered Costs
The price cut is not a sign of financial loss but reflects improved efficiency. The day before the price adjustment, DeepSeek began internal testing of the new V4.1 Flash model, which was announced to be more powerful and faster while using less computing power. Technological improvements have reduced the cost of generating the same content by a factor of several times. For example, if manufacturing a similar model previously cost 1 yuan, it now only costs 0.3 yuan, allowing DeepSeek to reduce prices without incurring losses.
4. Smart Pricing Based on Usage Patterns
DeepSeek continues to use a tiered pricing strategy, similar to how electricity prices vary during peak and off-peak hours. During business hours (9-12 AM and 2-6 PM), prices double to accommodate the high demand. This strategy is efficient because servers are fully utilized, and higher prices during these times generate more revenue. At off-peak times, prices are lower to attract developers and students who have more flexibility with their schedules, effectively using idle computing resources.
5. A Marking of a New Phase in Large-Model Pricing Competition
The previous price wars among manufacturers focused on offering extremely low prices or even free services, with the goal of attracting users at any cost. This new approach by DeepSeek demonstrates a shift towards a more sustainable model where products are priced differently based on their purpose and target audience. High-end models will continue to be expensive for large customers, while lower-end models will be priced competitively to attract a wider range of users. This will lead to more stable prices for AI services, with fewer instances of either extremely expensive or free but unreliable services.