Core Summary
This news article discusses the debate within the developer community regarding whether AI-generated code should require manual review. The opposing views of two prominent developers, Mitchell Hashimoto and Uncle Bob, have sparked significant discussion. Fundamentally, this reflects a transformation in the way code is produced in the AI era: whereas in the past, quality was ensured through manual line-by-line inspection, some now advocate for replacing this with a rigorous automated testing system. The approach suggests adjusting strategies based on the importance of the code—core code must be strictly controlled, while large amounts of “less critical” code can be generated and verified by AI, using a “layer of cheap code” to protect the “precious core code.”
The Conflict Between the Two Experts: Is Code Review a Fundamental Principle or an Old Habit?
- The “I Read Code” Camp: Mitchell believes that as long as one is responsible for the code (submitting, deploying, maintaining it), they must review AI-generated code as a professional necessity. Cindy, an expert in distributed systems, puts it more bluntly: “If you don’t understand how the code works, you can’t debug it; you don’t have the right to claim control over it—customers won’t trust a supplier that can’t even manage their own code.”
- The “I Don’t Read Code” Camp: Uncle Bob, who has been writing code for 60 years, insists on not reviewing AI-generated code at all. Instead, he relies on a set of checks: unit tests, natural language scenario tests (Gherkin), and mutation tests (intentionally altering the code to see if the tests fail). As long as the code passes these tests, he is confident in its quality.
The conflict arises from the question of whether to sacrifice efficiency for manual review or to replace it with a systematic validation process.
The “Vibe Slippery Slope”: Why Not Reading Code Could Lead to Uncontrollable Problems?
Open-source engineer Christine highlights the “Vibe Slippery Slope” phenomenon: initially, developers carefully reviewed AI-generated code; however, as AI becomes faster, their scrutiny slackens, leading to a tendency of “coding by feel.” This isn’t necessarily laziness but rather an inertia that is hard to break. Moreover, AI can generate hundreds of lines of code at once, which is beyond human review capacity—even experienced developers may not catch all the bugs in such a large amount of code. Over time, code quality deteriorates, and problems become difficult to fix.
Uncle Bob’s Smart Approach: Using “Cages” Instead of Constant Surveillance
Uncle Bob’s strategy doesn’t involve ignoring the issue but replacing manual monitoring with systematic constraints:
1. Have AI Write Tests: When AI generates code, it must also generate corresponding test cases (e.g., “Input 1+1, Output 2”).
2. Use AI to Write Verification Tools: AI can be used to create tools that check the reliability of these tests (e.g., whether edge cases are covered).
3. Multiple Layers of Validation: Use mutation tests and quality assurance processes to make it difficult for AI to cheat; altering tests would require modifying a complex network of related tests, which is costly.
4. Spot Checks on Critical Elements: He personally reviews the most critical natural language scenario tests (e.g., “Users should receive notifications after successful payment”) to ensure the overall direction is correct.
This approach has been implemented in the open-source tool old-coder, available for developers to use.
Code Has Different Levels of Importance: Not All Code Needs to Be Reviewed Line by Line
Theo, founder of t3.gg, categorizes code into four levels:
1. Garbage Level: One-time scripts (e.g., file organizers) that don’t require any review.
2. Annoying but Not Critical: Problems with these can be annoying but not catastrophic (e.g., personal blog plugins).
3. Job-Ending Level: Errors in these systems can lead to job loss (e.g., company core systems).
4. Lethal Level: Errors in these systems can be fatal (e.g., medical devices, aircraft control systems).
In the past, due to high code costs, only the latter two levels were closely monitored; now, with AI making the first two layers virtually free, large amounts of less critical code can be used to test the more important ones. For example, 1000 lines of test code can be generated for a single line of core code to perform stress tests and simulate edge cases—something previously unfeasible due to high costs.
The New Approach in the AI Era: Using Cheap Code to Protect Precious Code
AI has changed the cost structure of code, altering the logic of quality assurance:
- In the Past: High-cost code required careful review; writing 1000 lines of tests for a single line of core code was wasteful.
- Now: Since AI-generated code is free, a “test pyramid” can be built using cheap code—with a large base of automated tests and a small amount of core code at the top. This approach uses the less critical code to expose issues in the core code efficiently and securely.
For instance, 100,000 lines of test code can be generated for heart pacemaker software to simulate various extreme conditions (low temperature, high pressure, signal interference), ensuring the core code is flawless—a task that was previously impossible with manual review.
In Conclusion
The essence of this debate is not about whether to read code or not but about how to redefine “code quality” in the AI era: core code must be strictly controlled (whether through review or automated testing), while non-core code can be verified more efficiently by AI. After all, the value of AI is not to make us lazy but to allow us to focus on what truly matters.