虎嗅

"Have you noticed that the quota for Codex is becoming less and less effective?"

原文:有没有发现,Codex 的额度越来越不经用了

Summary of Key Points

This article reveals the harsh and hidden business logic behind current AI programming subscription services, using OpenAI Codex as an example: Although you pay a fixed monthly fee, the amount of computing power and working hours you actually get are being subtly reduced by the manufacturers through model upgrades and changes in billing policies.

In simple terms, previously, subscription software was like a buffet where you could eat as much as you wanted as long as you didn't waste it. Now, it's more like a limited-time, limited-quantity meal plan, where the restaurant owner (the AI manufacturer) can reduce the portion size at any time. Moreover, since new models require more computing power, what used to last a week might now only last half a day. Even worse, this reduction is not directly reflected in the bill but is hidden in complex algorithms and vague progress bars, causing users to spend more money without realizing it, and sometimes even face the situation where they can't buy enough resources even by paying extra.

---

In-Depth Analysis: Why Are AI Subscriptions Becoming Less Cost-Effective?

To help you fully understand the behind-the-scenes reasons, we break down this phenomenon into the following five aspects:

1. **Hidden Reductions**: Models Are Getting Better, but Your Money Is Getting Less Valuable

Many users notice that, despite the subscription tier remaining the same, projects that used to be completed with the available quota now only last half way. This is not an illusion but a result of computing power inflation caused by model evolution.

  • Previous Logic: Older models (such as early GPT-4 or Sol) had shorter thought processes and used fewer tools, working efficiently with less resource consumption.
  • Current Logic: New models (like GPT-5.6 Sol or Astra) are more sophisticated but also more verbose. To provide more accurate answers, they engage in longer reasoning chains and repeatedly use various tools for verification. This is like a worker changing from working directly to first drawing a plan, then finding materials, trying out the solution, and making adjustments, with each step consuming more of your “fuel” (tokens).
  • Conclusion: The official claim of a 18% increase in efficiency may only apply to individual tasks. However, if you consider the entire workflow, the increased resource consumption means that what used to last a week might now only last three days. This is why users feel that the quota has decreased—not because the number has changed, but because the value per unit of the quota has diminished.

2. The Mystery of Billing: The Percentages You See Are Actually the Manufacturer's “Black Box”

The article highlights a critical issue: users cannot accurately determine how much they are spending and can only see a vague progress bar.

  • The Deception of the Progress Bar: Manufacturers like OpenAI only display the percentage of usage without showing the actual number of tokens consumed or the equivalent cost. It's like a gas gauge at a station that only shows “half a tank left” without indicating the type of fuel (92 or 98 octane) or whether the tank has been secretly reduced in size.
  • The Rise of Third-Party Tools: Tools like NerfTrack are popular because users want to verify their suspicions through conversions. However, since the manufacturers do not disclose the specific billing formulas (e.g., the weight coefficients for different models and tools), users still cannot make accurate calculations.
  • The Power of Vagueness: This ambiguity in billing gives manufacturers significant control. When users complain about reduced quotas, they can argue that it's because the models are more resource-intensive, not that the quotas have been cut. This monopoly on defining the terms puts users at a disadvantage when seeking compensation.

3. The Business Calculus: The “Loss Threshold” Behind the Subscription Model and User Screening

Why do manufacturers do this? Because the cost of AI computing power is extremely high, while subscription fees are relatively fixed, creating a large profit margin or even the risk of loss.

  • SemiAnalysis’s Calculation: Research shows that OpenAI starts losing money when user utilization exceeds a certain threshold (e.g., the Plus tier at over 11.4%). This pricing model assumes that most users will not use the full capacity from the beginning.
  • Screening Heavy Users: For users who use AI intensively, such as authors writing code every day, the manufacturer is essentially “losing money while still generating traffic” or forcing them to upgrade to more expensive tiers.
  • Dynamic Adjustment Strategies:
  • Price Reduction Traps: A price cut on the Business tier in April seems beneficial but is accompanied by a reduction in the quota, potentially reducing the actual cost-effectiveness.
  • Purchase Restrictions and Suspensions: The suspension of the Pro 20X tier in September was claimed to be due to insufficient computing power, but the real purpose was to control costs and prevent high-value users from consuming too many resources.
  • Conclusion: Manufacturers are shifting the risk of high computing costs to users by dynamically adjusting quotas, limiting certain features, and suspending premium services, or by targeting users willing to pay more to maintain profitability.

4. Sudden Rule Changes: The Disruptive Shift from Unlimited Access to Pay-Per-Second Billing

The article compares AI subscriptions to mobile data plans and software subscriptions, highlighting the significant decline in user experience.

  • Certainty in Traditional Subscriptions:
  • Mobile Data: 10GB is 10GB; you run out, and the service stops, with clear rules.
  • Software Subscriptions: You pay for a year and get unlimited usage for 365 days, with stable rules.
  • WeChat Reading: The monthly fee is fixed, and you can read as many books as you want, with simple rules.
  • Uncertainty in AI Subscriptions:
  • The Fast Mode Trap: The speed is 1.5 times faster, but the resource consumption is 2.5 times higher. It's like paying more for faster highway traffic with no upper limit.
  • Frequent Rule Changes: Rules are changed frequently, such as removing the five-hour usage limit in July and reinstating it in August, or reducing quotas after a price cut in April. This lack of consistency makes it difficult for users to plan their work schedule.
  • Psychological Discomfort: Users are accustomed to the idea that a subscription means permanent access, but now they feel like they are in a “gaming” situation where they must constantly manage their resources and worry about running out of “tokens” (computing power), similar to managing a character’s health in a game.

5. The User Dilemma: Developers Who Are “Hostaged” by AI Manufacturers

Finally, the article discusses the real situation of developers today: They cannot do without AI but are also at the mercy of the AI manufacturers.

  • Technical Dependence: Many developers cannot write efficient code without AI tools. AI has evolved from an auxiliary tool to a core productivity tool, leaving them with little bargaining power when facing unreasonable manufacturer policies.
  • Resilient Strategies:
  • Multi-Platform Setup: Since one manufacturer is unreliable, developers use multiple alternatives. The author mentions setting up token plans for different models, choosing the one that suits each task to complete the work. This is a typical “decentralized” approach to survival.

Cost Comparison: Although all manufacturers are reducing resources, the cost of using AI is still much lower than hiring a human programmer. Therefore, users have to accept this exploitative situation and switch between manufacturers to find the most cost-effective combination.

Future Outlook: As AI capabilities become more widespread, this “subscription + dynamic quota” model may become the standard in the industry. Users will need to manage their AI budget like they manage a fluctuating currency, accepting “floating values” rather than fixed prices.

---

Recommendations for Ordinary Users

1. Don’t Just Look at the Monthly Fee: Before subscribing, try to understand the token consumption rate of the model for different tasks. If possible, use the free quota or a lower-tier subscription to test your typical workloads and see how much you will consume in a week.

2. Establish a Multi-Platform Backup: Don’t put all your eggs in one basket. Prepare AI services from at least 2-3 different manufacturers so you can switch quickly if one experiences quota issues or policy changes.

3. Follow Community Feedback: Platforms like the OpenAI community, tools like NerfTrack on GitHub, and developer forums are great sources for detecting hidden reductions. Official announcements are often delayed and vague, while user data from these communities is more accurate.

4. Adjust Your Expectations: Accept that AI subscriptions are not unlimited. Treat them like “pay-per-use cloud services” rather than traditional software licenses. Plan your tasks wisely to avoid major rework when quotas are low and complete essential work when quotas are sufficient.

In summary, the “subscription” model in the AI era is essentially a battle for computing power costs. Manufacturers aim for maximum profit, while users seek maximum efficiency. In this battle, staying clear-headed and flexible is the key to survival for developers.