Summary of Key Points
This article highlights a seemingly contradictory phenomenon in the use of AI: although the cost per individual AI inference unit (Token) is continuously decreasing, people are becoming more frugal when using advanced AI models, making them feel more expensive to utilize. The reason behind this is that the demand for AI is growing much faster than the supply of computing power. Model companies’ revenue has increased by 10 times, while computing power has only increased by 3 times. Platforms adjust resource allocation through price hikes and quotas; at the same time, users’ reliance on AI has deepened (using it for more complex tasks consumes more Tokens), resulting in an overall increase in spending rather than a decrease. Ultimately, the cost of computing power is borne by multiple parties, but not evenly. High-value tasks (such as coding and trading) get priority in resource allocation, while ordinary users have to compromise between cost, performance, and privacy, leading to a usage pattern where “advanced models are used sparingly, low-cost models are used freely, and local models are chosen for privacy protection.”
Detailed Analysis
1. Why Are Advanced AI Models More Expensive to Use? Higher Package Prices and Quotas Make People Hesitate
Many users struggle with decisions when using advanced models like Fable5: “Is this question worth solving with the most expensive model? Can the context be simplified? Should I try modifying it instead of running the model again?” In contrast, they are more casual when using cheaper models like MiniMax—asking whenever they need to. This isn’t about being stingy; rather, package prices have risen:
- The Coding Plan from iFlytek used to cost over 600 yuan per quarter but now has a monthly fee close to that of the previous quarter.
- New packages have adopted a points system, where using the strongest model during peak times requires three times as many points for deduction, with a limit of 400 uses in five hours and 2000 uses per week.
These changes force users to be more cautious when using advanced models, fearing they will exceed quotas or spend too much. Some users have even been banned for being too frugal (as mentioned by the author).
2. Why Could Computing Power Become More Expensive? Demand Is Growing 10 Times Faster Than Supply
Articles point out that model companies like Anthropic have seen their revenue increase by 10 times year-on-year, while computing power has only increased by 3 times. There are three possible ways to generate 10 times more revenue from the same amount of computing power:
- Model companies have higher profit margins (for example, Fable’s inference margin may exceed 80%).
- More computing power is being allocated for processing user requests rather than training.
- The cost of computing power itself has increased.
More importantly, cloud providers no longer price based on “cost plus profit.” For the same amount of electricity, users are willing to pay significantly more for using it to run AI that can generate revenue for businesses.
3. Even Though Tokens Are Cheaper, Why Isn’t the Total Spending Decreasing? Because You’re Doing More With AI
Although data shows that the cost of AI inference has decreased significantly over the years (for example, the median price for models with similar capabilities has dropped by 50 times), total spending may not have reduced. The reasons are:
- In the past, users might have asked the AI a single question; now, they use it for complex tasks that consume 5 to 30 times more Tokens than simple conversations.
- It’s like paper becoming cheaper, but if you start printing materials for an entire building, the total cost naturally increases.
4. Who Bears the Cost of Computing Power? Not Everyone Shares Equally
The increase in computing power costs is borne by chip companies, cloud providers, model companies, businesses, and individuals, but not evenly:
- Businesses: Divide routine tasks among cheaper models and use advanced models only for high-value tasks (like coding and trading), sometimes shifting the cost to downstream parties.
- Ordinary Users: Either pay more, wait in queues during peak times, or accept less powerful models.
- Platforms: Prioritize tasks that generate high revenue (such as enterprise-level code generation), while lower-value tasks like organizing old photos or telling stories to children may be limited due to difficulty in proving profitability (e.g., quotas and traffic restrictions).
5. Can Local Deployment Avoid High Prices? Freedom Comes at a Cost
Some people want to deploy AI locally to ensure privacy and independence from platform limitations. However, the barriers are high:
- High Hardware Costs: GPUs capable of running advanced models (such as the H100) are expensive, and additional costs are associated with power consumption and maintenance.
- Model Performance: Open-source models, although free, often fall short in performance compared to cutting-edge proprietary models like GPT-4.
Therefore, local deployment can only handle private or simple tasks and cannot replace advanced models. Users ultimately have to choose between using advanced models sparingly, low-cost models freely, or local models for privacy.
In essence, this article argues that while the basic capabilities of AI are becoming more affordable, the most advanced options are getting more expensive. Platforms use pricing and rules to make users pay for the specific capabilities they need, and ordinary users must balance cost, efficiency, and privacy. How much do you spend on AI each month? Which models do you use? Feel free to share your thoughts in the comments!