Summary of Key Points
As AI agents transition from being chat tools to performing actual business tasks—such as updating CRM data, processing refunds, and deploying code—the nature of their errors has evolved beyond mere nonsense. While the systems themselves may not report errors, subtle issues like booking incorrect flights, entering wrong customer information, or failing to process refunds can lead to real losses. Therefore, platforms like Google, Microsoft, and Grafana are implementing additional quality control measures for agents. These include using observability to record the entire process of an agent’s actions (similar to monitoring videos) and evaluation to determine whether the results are correct. The combination of these two approaches creates a closed-loop system that enables the discovery, correction, and optimization of errors, while also forcing companies to clarify their business standards and avoid becoming dependent on a single platform.
Why Quality Control is Suddenly Necessary for Agents That Are Doing Real Work?
Previously, agents were only used for chatting, and users could immediately correct any misunderstandings. Now, however, agents have the authority to modify data, make financial transactions, and send emails, meaning that their mistakes can have direct practical consequences. For example, when booking a flight, the system may return a “success” signal (with all technical indicators appearing normal), but the actual booking might be for the day after or the passenger’s name could be incorrect, resulting in wasted travel time for the user upon arrival at the airport. Moreover, as models are updated, prompts are revised, and new tools are integrated, agents’ behavior can change subtly, leading to errors that were not present before.
Traditional monitoring systems can only detect whether the system has crashed (such as interface timeouts or high CPU usage) but cannot verify whether the business tasks have been completed correctly. The case with Anthropic’s Claude Code is a typical example: users felt it had become less efficient, but the model itself had not degraded; the problem stemmed from reduced default reasoning intensity, bugs in context cleaning, and poorly optimized prompts—issues that traditional monitoring would not have detected.
Evaluation and Observability: One Determines Accuracy, the Other Tracks the Process
These two concepts are often confused, but they serve completely different purposes:
- Observability: It records what happened during an interaction—what requests were made by the user, which tools the agent used, what parameters were entered, how much time each step took, and which data was modified. It’s like installing a “dashcam” for the agent, allowing for a review of the process to identify the cause of errors (whether it was due to a misunderstanding in the model or incorrect parameter settings).
- Evaluation: It determines whether the task was completed correctly—has the refund been actually processed? Has the ticket order been canceled? Has the code passed testing? Have any sensitive information been leaked? For instance, if an agent says “refund processed,” evaluation would check the actual status in the payment system, not just rely on the agent’s response.
Microsoft puts it clearly: Observability tracks what happened, evaluation determines whether the task was done well, and optimization decides where improvements are needed. Without observability, it is impossible to identify the root cause of errors; with only observation, it is hard to distinguish between minor issues and major mistakes.
Quality Control Is More Than Just Adding a Monitoring Tool
Setting up quality control is not as simple as adding a monitoring panel; it requires establishing a closed-loop process:
1. Track Everything: Every action taken by the agent should be recorded in detail (which tools were used, which database fields were modified).
2. Identify Failures: Errors in business processes must be identified from the recorded data (such as booking incorrect flights or missing identity verifications).
3. Use These as Testing Cases: Real failure cases should be anonymized and used to create test scenarios for agents (for example, when booking a flight for tomorrow in Shanghai, the system should automatically check the date and passenger details).
4. Modify and Re-Test: The system should be updated based on these errors (e.g., by adjusting prompts or adding permission checks), and then re-tested with the modified test cases.
5. Roll Out Gradually: Start using these new measures with a small group of users to monitor for any new issues.
The combination of automated rules (such as checking if an order was successfully generated or if the amount is correct) and human judgment (for example, evaluating whether customer service responses are helpful) ensures efficiency and reliability. High-risk cases involving money or permissions must always be reviewed manually.
Quality Control Forces Companies to Define What “Success” Really Means
Many companies often say things like “customer service should be more helpful” or “salespeople need to understand customers better,” but these vague requirements are meaningless to agents. When implementing quality control, these goals must be translated into concrete, verifiable rules:
- “Customer service should be more helpful”: Does it mean resolving the user’s problem or just providing comfort? At what point does a human intervention become necessary?
- “Salespeople need to understand customers”: Data from one customer should not be mistakenly entered into another’s CRM.
- “High-quality reports”: Key conclusions must be based on reliable evidence and cannot be fabricated.
These rules are embedded in the details of business processes, such as compliance requirements for financial institutions or after-sales standards for e-commerce companies. Over time, they become valuable assets for a company. Large companies compete to see who can help them implement these rules effectively (Google wants to integrate them into its agent platforms, while Microsoft aims to incorporate them into its Foundry platform). To avoid being locked in to one provider, companies prefer using open standards like OpenTelemetry for tracking agent activities, so that they can transfer this expertise to new systems or platforms.
The Competition Among Agents Has Changed
In the past, AI was about who could answer questions most intelligently; now, it’s about which agents can perform tasks reliably. Actions are visible, results can be verified, and mistakes can be corrected. The goal is not for agents to never make errors, but for them to know when to stop or ask for help, and for the system to record their actions so that these can be used as learning opportunities. Quality control mechanisms are becoming the key to integrating agents into core business processes.