虎嗅

From a vague instruction to 131 chat screenshots: How I continuously corrected the AI

原文:从一句模糊指令到131张聊天截图,我是怎样不断纠正AI的

Summary of the Key Points

This article, which appears to be a casual user’s guide on training AI to process chat screenshots, actually demonstrates the entire process of using AI to complete complex, non-standard tasks through a highly demanding project: obtaining a series of continuous chat screenshots that can be used as evidence. It starts with a vague instruction to “organize the chat screenshots,” and after encountering several pitfalls and adding dozens of practical validation rules, 131 error-free screenshots are produced. This experiment shatters the myth that “writing a universal prompt will enable AI to do everything” and provides a practical work approach that can be applied both by individuals and businesses.

---

Detailed Analysis

90% of people fail to use AI effectively because they define the “acceptance criteria” too vaguely

When people encounter problems with AI, their first reaction is often that “AI is too stupid.” However, the author’s first round of mistakes clearly illustrates this point: he initially asked the AI to come up with a plan before starting the task, but the quality report submitted by the AI stated that “the number of untranscribed voices was 0,” even though half of the screenshots contained untranscribed audio. This is not because the AI is deliberately misleading; rather, the definition of “qualified” is completely different for both of them. To you, a “qualified screenshot” means that all audio has been transcribed into text, the sections are coherent, and it can be used as evidence. For the AI, “qualified” simply means that there are no obvious blank transcription areas. This is similar to the problems that arise when you outsource app development or assign tasks to subordinates: you ask for a “good-looking and user-friendly page,” but what they deliver often falls far short of your expectations. Vague requirements cannot lead to precise results; if you don’t specify your needs clearly, the AI will handle the task in the most efficient way possible, which may not meet your expectations.

AI has strong execution capabilities, but it cannot “understand” your intentions; you need to break down the tasks into clear, rule-based steps

In the second round of the experiment, the author transformed all vague requirements into concrete, objective validation criteria that the AI could easily check. For example, to determine if the audio has been transcribed correctly, you can’t just check for blank transcription areas; instead, you need to verify three conditions: 1) no blank or loading transcription areas, 2) text actually appears below the audio bubbles, and 3) no circular animation in the audio area. Similarly, to determine if two screenshots are consecutive, you need to find common messages, images, or timestamps between them. AI is highly efficient, but it lacks human understanding; by breaking down your requirements into actionable checks, you can have it verify tens of thousands of screenshots in seconds with greater accuracy than 10 interns could achieve in hours. Many companies spend millions on large models but fail to use them effectively because they don’t do this step of defining the rules and expect the AI to figure it out on its own, resulting in unusable output.

For tasks involving “authenticity,” no room for AI to “improvise”

The fourth round of mistakes in the experiment is particularly enlightening: when WeChat starts playing long audio, it may skip frames, and the source video doesn’t contain the complete sequence of frames before and after the skip. The AI, trying to create a continuous screenshot based on the conversation’s logic, was stopped by the author’s rule that it couldn’t fabricate missing parts; it had to either mark the gaps or split the screenshot into segments. This highlights a major weakness of large models, which tend to “complete the logic” automatically, leading to inaccuracies. Many people use AI for financial report organization and evidence analysis, but they often grant the AI the authority to make reasonable inferences based on common sense. If you don’t specify that all content must come from the original sources and that any fabrication is prohibited, the AI may create false records or data, which can be problematic when used as evidence or for public release.

There is no such thing as a “universal prompt”; an efficient AI workflow involves gradual iteration and rule development

In the final analysis, the author noted that even if the AI were to come up with a plan first, it would still be impossible to anticipate all possible scenarios in advance. Details like the automatic scrolling of previous chat content at the end of a screenshot or the unannounced skipping of long audio during playback are impossible to predict. The so-called “universal prompts” circulating online are essentially a waste of time. An effective AI workflow involves continuous iteration: starting with a small sample of data, identifying and fixing issues, and developing a set of validation rules (in this case, 43 regression tests). Each new issue leads to the addition of another rule, ensuring that the same mistakes are not repeated when processing larger datasets. This approach is consistent with the implementation of AI in all industries; no AI technology can significantly reduce costs from the start; useful tools are developed through repeated processing of real data, addressing previous challenges.