The Orchestrator's Fallacy: Why You Still Need to Understand the Code
Reading AI-generated code takes more skill than writing it, and when an agent gets something wrong, fixing the pull request is only the start.
Executive Summary: There's a popular idea, captured in the phrase "vibe coding" [1], that as AI agents take over code generation, engineers will become prompt writers who don't need to read the code anymore. The research says the opposite. In one controlled trial, experienced developers believed AI made them faster when it actually slowed them down [3], two-thirds of developers say they've run into AI output that is almost right, but not quite [4], and in one large study nearly half of AI-generated code samples had a security flaw [5]. Catching those problems takes more knowledge of the system than writing the code did. And when an agent gets something wrong, fixing the pull request is the easy half. The harder half is changing what the agent works from (its context, its checks, its templates) so the mistake is less likely to come back.
The Code-Blind Orchestrator
You've probably heard the pitch: engineers describe what they want, agents write it, and nobody has to read the code closely again. The idea took off in February 2025, when OpenAI co-founder Andrej Karpathy coined the term "vibe coding" for a way of working where you go with the flow and forget the code is even there [1]. He was describing his own quick, casual projects, but the phrase became shorthand for prompting and shipping without reading the code. Tellingly, he has since moved on. By 2026 he was drawing a clear line between vibe coding and what he calls agentic engineering, where the engineer does not blindly accept what the agent generates [2].
The best evidence we have doesn't back it up. In a randomized controlled trial, METR gave 16 experienced open-source developers 246 real tasks in their own mature repositories, randomly allowing or disallowing AI tools for each one. With AI allowed, tasks took 19% longer, yet the developers believed AI had made them about 20% faster [3]. METR is clear that this is one setting with early-2025 tools, and newer tools may well do better. It's still a useful warning: on a large, real codebase, working with AI output carried real overhead, and even experts couldn't feel how much it was costing them.
Developers describe the cost directly. In Stack Overflow's 2025 survey, the most common frustration with AI tools, cited by 66% of respondents, was solutions that are almost right but not quite. The second most common, at 45%, was that debugging AI-generated code is more time-consuming. More developers distrust the accuracy of AI output (46%) than trust it (33%), and experienced developers are the most cautious group [4].
Some of what gets missed is serious. Veracode tested code from more than 100 large language models across 80 coding tasks and found that 45% of samples introduced a known security flaw, with newer and larger models doing no better [5]. That code compiles and reads cleanly, so catching it takes automated scanning plus a reviewer who knows what the code was supposed to do.
Reading Is Harder Than Writing
Part 1 described AI output as locally correct but globally wrong: each change looks right on its own but was written without the context your senior engineers carry. That is exactly what makes it hard to review.
When you write code yourself, you build a mental model of the system as you go. When you review generated code, you have to rebuild that model from the outside and then check the code against it. Mostly you're looking for what isn't there. In some cases, AI-generated code handles each state change correctly on its own but never makes those operations idempotent. When the same request arrives more than once, the state change runs again and breaks the system. The fix is changing those operations, so a repeated request leaves the system exactly where it was. None of that shows up in a clean diff unless you already know how the system behaves.
That's why I think the code-blind orchestrator idea has it backwards. Directing agents well takes more technical depth than writing the code yourself. Good specs and prompts matter, and they're worth getting better at, but they sit on top of engineering judgment. If you can't trace how data moves through the system, you can't tell good output from almost-right output, and almost-right is what makes it to production.
Guardrails Were Never About AI
None of the safeguards that help here are new. Good engineering teams have long relied on automated checks so people don't have to catch everything by eye. The architectural version is what Neal Ford, Rebecca Parsons, and their co-authors call fitness functions: automated tests that give an objective read on whether the system still has the qualities you care about. In a 2025 conversation about AI, Ford described writing architecture rules as pseudocode and having generative AI turn them into fitness functions that run in the build pipeline [6].
What AI changes is the speed at which their absence hurts. When people write code slowly, duplication and drift build up slowly. When agents produce changes far faster than people can read them, the same gaps compound much faster. GitClear's analysis of 623 million code changes shows the trend: copy/pasted code rose to 15.7% of changed lines while refactored code fell to 3.8%, down from 21% in 2022 [7]. As Part 1 noted, that's a correlation with rising AI use, not proof of cause.
The gates worth having are the same ones you would want for any team:
- Contract and schema checks at API and data boundaries, so a change that breaks an agreement between services fails before anyone reviews it.
- Architecture checks for dependency and layering rules, plus a duplication threshold, so drift fails the build instead of waiting for someone to notice it.
- Security scanning on every change. Given the Veracode numbers, this one isn't optional [5].
None of these should care who wrote the change. A junior engineer, a senior engineer, and an agent all go through the same gates.
Thoughtworks lands in a similar place in its 2026 Looking Glass report: humans and AI should build software together, engineers stay responsible for architectural integrity, and governance belongs in the delivery pipeline rather than in someone's head. It also puts the typical delivery gain it sees from AI at up to 15%, far below what these tools are usually sold on [8]. Productivity numbers vary widely by study and setting, which is itself the point: the gains depend on the system they land in.
Fix the PR, Then Fix the System
When a junior engineer introduced a messy pattern, a senior engineer often fixed it in review, explained why, and the junior did better next time. That loop does not work with an agent. As Part 1 noted, a model starts each session knowing only what it is given. Correct it in a review comment and it is likely to make the same mistake tomorrow.
The pull request still has to be fixed before it ships. The more useful question comes after that: why did the agent do it? In my experience the answer usually falls into one of these buckets.
- It didn't know. The agent had no way of knowing an internal service already existed, or what a domain rule was. Put it where the agent reads: a decision record, an instruction file, a catalog of existing services.
- Nothing stopped it. The rule was known but never enforced. Turn it into a test or fitness function so the build fails next time.
- It keeps happening. The same kind of mistake shows up across changes, which usually means the template or golden path the agent starts from is wrong. DORA found a direct link between the quality of an organization's internal platform and how much value it gets from AI [9].
Sometimes it's simpler than that: the spec left a gap and the agent filled it. Tighten the spec, and keep the work small enough to review in one sitting.
A caveat: models are probabilistic, so better instructions lower the odds of a repeat without ruling it out. That's why I want both, and why the test matters more. The instruction makes the mistake rarer, and the test catches it when it happens anyway.
This is Part 2 of a 3-part series on the operational realities of AI-native engineering. Part 1: The Verification Tax: Faster AI Code, Less Stable Delivery. Next up in Part 3: The 3-Person AI Pod: How to Save the Junior Engineer.
References
- Andrej Karpathy, post on X introducing "vibe coding" (February 2, 2025): the post that coined the term.
- Andrej Karpathy, "Sequoia Ascent 2026: Software 3.0, Agentic Engineering, and Jagged Intelligence" (April 2026): Karpathy's own summary of his Sequoia AI Ascent fireside chat.
- Becker, Rush, Barnes & Rein, "Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity," METR (July 2025): randomized controlled trial, 16 developers, 246 tasks.
- Stack Overflow, "2025 Developer Survey: AI" (2025): about 49,000 respondents.
- Veracode, "2025 GenAI Code Security Report" (2025): 100+ large language models, 80 coding tasks.
- Thoughtworks Technology Podcast, "How fitness functions can help us govern and measure AI," with Neal Ford and Rebecca Parsons (March 2025): two of the authors of Building Evolutionary Architectures on applying fitness functions to AI.
- GitClear, "The Maintainability Gap: 2026 AI Code Quality Research" (January 2026): 623 million code changes, 2023–2026.
- Thoughtworks, "Looking Glass 2026: AI and Software Delivery" (January 2026): Thoughtworks' annual technology trends report.
- Google Cloud DORA, "State of AI-assisted Software Development" (2025): survey of nearly 5,000 technology professionals.