Summary of Key Points
This article focuses on the “skills” of large-scale models (which can be understood as rules, processes, or system prompts provided to AI). The main finding is that while adding skills to AI theoretically improves task performance, it actually leads to a phenomenon known as the “regression tax”—newly solved problems are offset by a decline in existing capabilities. The article explains the causes of the regression tax through experimental data and technical principles and offers engineering solutions.
Why Does Longer Skill Descriptions Lead to Unstable Performance?
Many people follow this process when writing skills for AI: initial version → trial run to identify issues → add constraints → further iteration → an increasing number of prohibited terms, resulting in skill descriptions growing from 400 lines to 2000 lines, yet performance remains inconsistent.
This is due to the “attention mechanism” of large-scale models. A Stanford paper titled “Lost in the Middle” demonstrates that AI’s attention to input content follows a U-shaped curve—it remembers the beginning and end well but forgets the middle. For example, if the instructions at the beginning state “be accurate” and the end states “do not omit anything,” the AI may overlook the 1000 lines of rules in between. As a result, more rules lead to longer documents, which in turn cause the middle content to be ignored, leading to more rule violations and creating a vicious cycle.
What is the “Regression Tax”? — The Cost of Enhancing New Abilities
Traditional evaluations only look at the overall improvement after adding skills but overlook an important factor: the increased score could come from either “getting 5 new questions right and all old questions right” or “getting 20 new questions right and 15 old questions wrong.” The latter scenario is much less reliable, and the cost of this “degradation of existing abilities” is what constitutes the regression tax.
The paper tested three sets of models and skill libraries with 486 office tasks (such as reading financial documents and editing Excel). The results showed that while adding skills led to 553 additional successes, it also resulted in 324 failures in previous tasks. These failures offset 59% of the total gains, leaving only a net increase of 229 successes. All 18 skill conditions exhibited regression, with none being able to enhance abilities without damaging existing ones.
The Two Main Causes of the Regression Tax: Penetration Effect and Misunderstanding Instructions
1. Penetration Effect: Unused Skill Descriptions Affect Decision-Making
Skill descriptions typically consist of two parts: the “description” (what the skill solves) and the “body” (the specific steps). Many AI frameworks embed the description in the context to help the AI determine whether to activate the skill. However, even if the body is not used, the concepts within the description can still influence the AI’s decision-making process.
For instance, if a skill describes handling sales data and the user asks for a report on monthly performance, the AI might automatically use the sales data-related skill, ignoring that the user actually needs a comprehensive report including other departments’ performance. This is the “penetration effect,” where the AI’s judgment is skewed by the skill description.
2. Misunderstanding Instructions: Following the Process Without Comprehending User Needs
This phenomenon is referred to as “grounding displacement” in the paper, meaning the AI fails to understand the user’s requirements. For example, if the user asks for the difference in public engineering expenditures between 1934 and 1946, the AI might use the wrong tables, years, or statistical methods, leading to incorrect results despite following the correct process.
Another issue is “verification displacement”: the AI completes the task according to the skill instructions without checking the results. For instance, an Excel formula might be correct, but the AI submits it without verification, or the scoring system may not recognize certain functions, resulting in incorrect evaluations.
How to Reduce the Regression Tax? — Three Practical Engineering Approaches
Drawing on software engineering practices, the article proposes the following solutions:
1. Hierarchical Architecture + Positive Instructions: Divide skills into layers (e.g., place core rules at the beginning or end and secondary rules in the middle), and use “only do Y” instead of “do not do X” (positive instructions are more likely to be followed by the model).
2. Structured Format: Use clear formats like markdown tables and bullet points to reduce the AI’s understanding effort (for example, listing rules in a table makes them easier to remember than long paragraphs).
3. Regression Testing + Gradual Deployment: Before adding new skills, test whether existing tasks are affected negatively; release new skills gradually (in a phased manner) and revert them if necessary, similar to software updates, to avoid issues due to sudden full-scale implementation.
In Conclusion
Adding skills to AI is not about more being better; it’s about adding the right rules precisely and controlling the regression tax. Just like managing software projects, it’s essential to enhance capabilities while preserving existing strengths.