Writing · Part 1 of 3

The Verification Tax: Faster AI Code, Less Stable Delivery

AI is helping teams write more code than ever. Much of that gain is being spent checking it.

Executive Summary: AI coding tools make individual developers faster, but that speed is not reliably reaching delivery. The clearest signal in the research is instability: pull requests are bigger, teams have far more code to review, and releases break more often. The cause is code that is locally correct but globally wrong. Each change looks right on its own, but it was written without the architectural and domain context your senior engineers carry in their heads. The organizations getting value from AI are the ones that give it that context and keep changes small enough to review.

The Productivity Paradox

The individual gains are real. In field experiments with nearly 5,000 developers at Microsoft, Accenture, and a Fortune 100 company, developers given an AI coding assistant completed about 26% more tasks [1].

What happens next is less encouraging. Faros AI's 2025 telemetry from more than 10,000 developers found that teams with heavy AI use merged 98% more pull requests, but those pull requests were 154% larger on average and took 91% longer to review. Across companies, Faros found no significant link between AI adoption and better delivery performance [2].

Google's 2025 DORA report, based on a survey of nearly 5,000 technology professionals, is more positive about speed: unlike in 2024, higher AI adoption was associated with higher delivery throughput. On stability it agrees with Faros. In DORA's words, AI adoption "not only fails to fix instability, it is currently associated with increasing instability." The report's central conclusion is that AI amplifies whatever system it lands in, and that the greatest returns come "not from the tools themselves, but from a strategic focus on the underlying organizational system" [3].

So the research disagrees on speed and agrees on stability. Teams are producing much more code, releases are getting less stable, and a large share of the time AI saves is being spent checking its output. That cost is the verification tax.

We Have Seen This Before

Engineering leaders have seen this pattern before in the junior engineer who copies code from Stack Overflow without understanding the system it is going into. That engineer solves the local problem, writes a new utility instead of finding the existing one, and ships logic that works on the happy path but breaks team conventions.

AI assistants and agents working without context behave the same way. A junior engineer eventually learns the system, though. A model starts each session knowing only what it is given, so the context has to be supplied deliberately, every time.

What the Code Data Shows

GitClear's January 2026 analysis of 623 million code changes found copy/pasted code rising to 15.7% of changed lines, while refactored code fell to 3.8%, down from 21% in 2022 [4]. GitClear measures these trends across repositories rather than tagging individual AI-written commits, so the data shows a correlation with rising AI use rather than proof of cause.

Developers see something different. In DORA's survey, 59% of respondents say AI has improved the quality of their code [3]. Both findings can hold at once. Each change looks better to the person who wrote it, while duplication grows and refactoring disappears across the codebase.

The Verification Tax

When code is produced faster than people can review it, review becomes the constraint. DORA describes how the work changes for individual developers: it "shifts from manual grind to deciding and verifying," including "assessing code that looks remarkably similar to correct code" [3].

AI-generated code is hard to review for a specific reason. Human mistakes tend to come with visible signs, such as messy formatting or half-finished logic. AI output reads cleanly, so the problems are in what it leaves out: an undocumented business rule, a security constraint, an existing service it should have reused. Reviewers have to check code that looks correct against knowledge that was never written down, and there is now far more of it to check.

The load also spreads beyond the reviewer. When every team ships more changes, every neighboring team has more to absorb. DORA ties AI-era instability to the same pattern: more change volume without updated guardrails and "golden paths" raises verification and coordination costs, and teams that depend heavily on other teams stay unstable even after adopting AI [3].

Context and Guardrails

When junior engineers leaned on Stack Overflow, good teams responded with context, patterns, and guardrails. AI needs the same, and the need grows as teams move from autocomplete to agents that change dozens of files in a single session.

Conway's Law says systems mirror the communication structures of the organizations that build them. For an AI agent, the communication structure is the context it can read. If domain boundaries and API contracts are written down, agents can follow them. If they live only in people's heads, agents cannot see them, and the code drifts.

  1. Engineer the context. Put architecture decisions, domain rules, API contracts, and conventions where the AI can read them: decision records, agent instruction files, and contract-first specifications. DORA found that AI's positive effect on individual effectiveness and code quality is amplified when AI tools can reach internal company data [3].
  2. Keep changes small. Break work into spec-sized units a person can review in one sitting. Working in small batches is one of seven capabilities DORA found make AI more effective, and teams that do it see AI's positive effect on product performance amplified [3].
  3. Automate the first review. Tests, static analysis, security scans, and architecture conformance checks should catch mechanical problems before a person looks, so human reviewers can focus on design and domain correctness. AI-assisted review is a useful addition to those checks rather than a substitute for them.
  4. Treat agents as platform customers. Platform teams should give agents what they already give developers: governed access to models, shared context, and golden paths with clear boundaries. DORA found a direct correlation between the quality of an organization's internal platform and its ability to get value from AI [3].

What This Looks Like in Practice

I have seen these practices work firsthand. A fellow engineer designed an agentic development framework around them and used it to build a platform with backend services, a frontend, and a generated client SDK. Every significant architecture decision is written down as a decision record. Agents start each session from instruction files and a catalog of what already exists. Services are built contract-first, so the API specification comes before the code that implements it.

In my view, two problems can change the most. Before that structure is in place, agents can regularly drift from architectural decisions that have already been made, and changes can arrive in batches too large to review with confidence. Once decisions are recorded where agents read them, and work is broken into spec-sized units, drift should become the exception, and review time can go to design questions instead of cleanup. This is my own assessment rather than a controlled study, which is why I am testing it in public.

Speed without architectural alignment is not productivity. It is just faster technical debt.

This is Part 1 of a 3-part series on the operational realities of AI-native engineering. Next up in Part 2: The Orchestrator's Fallacy: Why you still need to understand the code.

References

  1. Cui, Demirer, Jaffe, Musolff, Peng & Salz, "The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers," Management Science (2025)
  2. Faros AI, "The AI Productivity Paradox Research Report" (2025): telemetry from 10,000+ developers across 1,255 teams.
  3. Google Cloud DORA, "State of AI-assisted Software Development" (2025): survey of nearly 5,000 technology professionals.
  4. GitClear, "The Maintainability Gap: 2026 AI Code Quality Research" (January 2026): 623 million code changes, 2023–2026.