虎嗅

"Insufficient Skills: Agents Need to Develop 'Production Systems'"

原文:Skill 不够了,Agent 需要拼“生产系统”

Summary of Key Points

This article focuses on two crucial concepts related to AI Agents: Skill and Harness. Skill involves encapsulating industry experience and work processes into “instructions” (such as code guidelines or business SOPs) to tell the Agent what to do. However, having these instructions alone is not enough; agents may be lazy, skip steps, or make decisions on their own.Harness, on the other hand, provides a “reliable workspace” for the agent by specifying which tools it can use, which actions require approval, how to verify results after completion, and how to backtrack in case of errors, transforming the agent from something that can “talk” into something that can perform tasks reliably.

As agents begin to handle longer and more important tasks (such as code modification or financial processing), the importance of Harness becomes increasingly evident. It determines whether the agent can evolve from a demonstration tool into a work system that businesses are willing to use.

I. Skill: The “Instruction Manual” for Experience, but Ineffective in Controlling Agent Behavior

Why are Skills so popular? Large models cannot retain all the experience of a company (e.g., code guidelines or compliance requirements). By breaking down this experience into individual task packages (e.g., “code review Skill” or “financial reconciliation Skill”) and providing them to agents when needed, it’s more efficient than giving them an entire employee handbook. However, Skills have a significant drawback: they are merely “soft constraints.” For example, even if you specify that “a backup must be taken and testing must be performed before modifying the database,” the agent might skip testing due to thinking the change is minor or claim it’s completed halfway through. This is similar to providing a new employee with an SOP; they may find it cumbersome and choose not to follow it, especially for longer tasks where missing a step can lead to widespread errors that are difficult to detect.

II. Harness: The “Workstation” That Keeps Agents on the Right Track

Harness is not a more advanced form of Skill but rather an environment that ensures agents act in a structured manner, similar to providing a workstation with safety barriers:

  • Material Area: Agents can access relevant documents and historical task records to know how to continue if they pause mid-way (solving the problem of forgetting).
  • Tool Area: Does it allow access to browsers or databases? Can the results of these actions be viewed after use (e.g., checking logs after code modification)?
  • Safety Barriers: Certain actions (like modifying a production database) require double confirmation; without proper checks, the next step cannot be proceeded with (these are hard constraints, not just reminders).
  • Verification Area: Results must be verified after completion (e.g., running tests after code modifications or checking records after data entry); errors can be rolled back if they occur.

For instance, while Skill instructs “back up → test → execute the migration,” Harness turns these into mandatory steps. Without backup records, the next button will be disabled; if the test fails, the agent cannot proceed. Experiments show that the proportion of agents completing the steps correctly increased from 56% to 86% with the use of Harness, highlighting the difference between relying on “hope” and having a system that enforces these rules.

III. Why Is Everyone Suddenly Talking About Harness in 2026?

Previously, Harness was considered an unremarkable component, but now it has become a core competitive advantage for several reasons:

1. Agents Are Taking On More Complex Tasks: They are moving from simple tasks like writing code to completing entire projects that involve multiple systems and steps, requiring proper state management (e.g., being able to continue work if another agent takes over midway) and error tracking.

2. Skills Have Become Standardized: As Skills adopt a standardized format, the ability to support a wide range of Skills has become a basic requirement, shifting the focus to Harness (for example, the same Skill may perform poorly in differentHarness environments).

3. Businesses Are More Willing to Use Agents: With agents moving into production environments, businesses need to have visibility into their actions, who authorized them, and who can be held accountable for any errors. Harness provides the necessary permissions, logging, and approval mechanisms to address these issues.

For example, OpenAI’s use of Codex initially slowed progress due to an unclear environment, but optimizingHarness (by providing agents with a dedicated workspace and access to logs/screenshots) accelerated development.

IV. Future Value: Vertical Harnesses Are More Valuable Than General Skills

While Skills are easy to replicate (e.g., a “legal search Skill” can be created by anyone), vertical Harnesses are much more challenging to develop:

  • For legal agents, Harness must know what to look for in different jurisdictions, when to stop due to conflicts of interest, and how to verify references.
  • For financial agents, it must manage exceptions during reconciliations, who has the authority to approve payments, and how to keep records.

Companies are not buying “intelligent agents” so much as agents with clear boundaries: they need agents that know what they cannot do, have evidence of their actions, and can be retracted if necessary. Therefore, the future lies in vertical Harnesses that understand the specific workings of industries and enterprise-level Agent Runtime systems (platforms that unify permissions, logging, and strategies).

V.Harnesses Are Not Universal: Avoid Overcomplicating Them

However, Harnesses also have their limitations:

  • Overcomplexity: Adding too many features (multi-agent orchestration, retries, sandboxes) can lead to unclear issues when something goes wrong.
  • High Cost: Multiple validation steps and retries consume more tokens and computing resources.
  • Overfitting: A Harness optimized for one task may fail when used for another.
  • Security Risks: Skills may contain scripts, and their sources and permissions need to be carefully verified during installation.

Therefore,Harnesses should be tailored to specific needs: use simpler versions for tasks like sending birthday wishes (no need for complex verification), more robust versions for online service modifications (with testing and rollback capabilities), and even stricter versions for financial transactions (with approval and auditing processes).

In the end, no matter how advanced the models become, Harnesses will remain essential. Just as intelligent employees require permissions and approval processes within a company,Harnesses serve as the organizational “management rules” for agents, enabling businesses to confidently delegate tasks.

In essence, the article emphasizes that the key to transforming AI Agents from capable communicators into reliable task performers lies in providing them with a set of “reliable rules”—this is the true value of Harness.