Est.

Reviewing Two-Thousand-Line Agent Diffs

Agents leave no trace of their reasoning, making architectural failures invisible to review.

Contributing Editor · · 10 min read
Cover illustration for “Reviewing Two-Thousand-Line Agent Diffs”
Agent-Generated Code · October 5, 2026 · 10 min read · 2,211 words

A two-thousand-line pull request opened by an autonomous coding agent is a different artifact entirely from the diffs that code review was designed around, one produced by a different process, and reading it line by line will not tell you what was actually built. The volume of agent-generated code has already exceeded what the review process built for human pull requests can absorb, and the rest of this piece is about what that means for how engineering teams verify what agents do to their codebases.

Why agent diffs have outgrown diff review

Coding agents do not work the way the tools that diff review was designed for assumed developers worked. An agent interprets a task, plans an approach, generates code, runs tests, and iterates, all inside a single closed loop that can touch dozens of files before a human ever sees a line of it. Traditional review tooling was built around a single prompt-response exchange: a developer writes a change, opens a pull request, and a reviewer reads it. Agents don't submit that kind of artifact. They submit the output of an entire autonomous work session, compressed into one diff.

The scale backs this up. The AIDev dataset puts a sharper point on what that growth looks like in practice: OpenAI Codex alone opened more than 400,000 pull requests in open-source repositories within two months of its release, a volume that makes "agents contributing thousands of PRs daily" a literal description.

Diff size has grown alongside volume. Review time has risen in step, yet a growing share of pull requests merge with no recorded review. Capacity collapse looks like a relaxation of standards, but it is what happens when the volume of work outstrips the humans available to check it.

Every number in this section describes a quantity problem: more commits, bigger diffs, less review time per line. But quantity is not what makes these pull requests hard to review. A diff twice as long as last year's diff is still something a reviewer could, in principle, read twice as long to get through. What agent diffs actually break is something else.

Structural differences between a two-thousand-line agent diff and a large human diff

A two-thousand-line diff written by an agent is the output of a different kind of reasoning process than a human would have used, and that process leaves almost no trace in the artifact a reviewer ends up looking at.

When a human engineer submits a large change, a reviewer has more to go on than the diff itself. An agent's diff carries none of that. Its intermediate reasoning, the tradeoffs it weighed, the alternative structures it considered and discarded, all of that existed somewhere in the agent's working process and none of it survives into the final diff. What a reviewer sees is the output, stripped of the decision path that produced it.

Agent diffs also mix in changes no human reviewer would expect to find bundled with a functional change. Separating what matters from what doesn't becomes a line-by-line forensic exercise, and it is one that gets harder, not easier, as the diff grows.

The deepest version of this problem is architectural. Consider a codebase like VS Code, which enforces a layered structure where foundational utilities belong in a specific base layer rather than wherever their first caller happens to live. The same is true of a more deliberate shortcut: an agent that routes a database query directly through a controller to shave latency, bypassing a prescribed service layer, is making an architectural decision. That decision's rationale exists only in a reasoning trace the agent has already discarded by the time the pull request is open. The diff shows the query. It does not show the layer it was supposed to go through.

How AI-on-AI review fails to solve the structural problem

The obvious response to an unreadable diff is to hand it to another AI and let that system do the reading. Routing a large agent diff through an AI code review agent processes the same artifact faster. It does not produce an understanding of what the agent actually built, because the second AI is working from exactly the same incomplete record as the first reviewer would.

Reviewers appear to sense that something about these diffs resists reading even when they don't name the problem explicitly. Human-AI Synergy in Agentic Code Review found that human reviewers exchange 11.8% more rounds of back-and-forth when reviewing AI-generated code than when reviewing human-written code, a gap that suggests the difficulty isn't just a matter of more lines to get through. A two-thousand-line pull request does not get a two-thousand-line review in practice, from a human or from a model standing in for one.

The data on code review agents operating without human oversight makes the limitation concrete. From Industry Claims to Empirical Reality found that pull requests reviewed only by a code review agent, with no human reviewer involved, merge at a rate 23.17 percentage points lower than pull requests reviewed by humans, and that the majority of closed, CRA-only pull requests fall into the lowest signal range the study measured. The reviewing AI is working from the same incomplete artifact as everyone else and drawing conclusions from it that often don't hold up.

Even the fixes that do get adopted carry a cost: Human-AI Synergy found that when agent suggestions are adopted into a codebase, they produce significantly larger increases in code complexity and code size than suggestions coming from human reviewers, so the review process itself can add to the structural burden it was meant to reduce.

None of this is an argument against tools like CodeRabbit, which catch real defects faster than no review would and genuinely help teams process volume they could not otherwise get through. The argument is narrower and harder to dismiss: even perfect defect-catching speed does not touch the category of failure that threatens a production codebase most, which is architectural drift, violation of structural constraints, and incoherence spreading across a surface an agent touched in a single session. That category of failure is what diff-level review, whether the reviewer is a person or a model, is worst equipped to catch, because the diff was never built to show it.

What agent diffs do to a codebase's architecture

The damage a large agent diff can do rarely appears as a bug inside a single function. It appears instead as a violation of the structural rules that keep a whole system coherent, and that class of violation does not register on a line-by-line read no matter how carefully the lines are read.

Agents optimize for finishing the task in front of them. They are not optimizing for the long-term coherence of the system they're working inside, and they often can't: they lack the kind of accumulated, whole-codebase familiarity that lets an experienced engineer sense when a change doesn't belong somewhere, and on large repositories an agent may not even hold the full codebase in its working context at the moment it makes a decision. The VS Code example is one illustration: a utility function that belongs in the base layer gets placed next to its first caller instead, a constraint violation no diff reader will catch unless they already knew the rule going in. The controller-level database query is a second illustration, a deliberate architectural shortcut whose justification lived only in a reasoning trace the agent never wrote down anywhere a reviewer could find it.

What reviewers actually spend their attention on confirms the gap. Understanding Dominant Themes in Reviewing Agentic AI-authored Code found that review comments on agent-authored pull requests cluster overwhelmingly around documentation gaps, refactoring suggestions, and styling or formatting issues, the concerns most visible on the surface of a diff, while testing and security concerns receive comparatively little attention. That pattern is consistent with reviewers responding to what the diff makes easy to see rather than to what actually matters most to the system's integrity.

That mismatch raises outcomes-level costs, not just comments. A telemetry study of a large developer population by Faros AI found that teams with high AI adoption merged substantially more pull requests and completed more tasks, and bug counts still rose by 9% across that same population. Individual velocity went up. System-level quality did not follow, because architectural drift accumulates underneath a stream of individually reasonable changes until it surfaces as a system-level failure.

None of this is a claim that agents write bad code at the level of a single function. The code an agent produces can be locally correct, pass its tests, and satisfy the immediate task, while still being wrong at the level of the system it was dropped into. That distinction is the entire argument.

Diagram: Review Comments Cluster on the Wrong Things. Visualizes: A ranked breakdown showing where human reviewers actually focus their attention on agent-authored pull requests versus what matters most to system integrity.

Why earlier enforcement only treats a symptom

The other dominant response to unreviewable diffs is to move enforcement earlier, catching violations before a human or an AI has to read the diff. This is a real improvement over catching problems at review time, and it is still not sufficient on its own.

The reasoning behind pre-commit and pre-merge gates holds up. A check that fails deterministically and blocks a merge is worth more than a review a tired developer rubber-stamps an hour into a long session. Traditional CI/CD placed its gates after code was already written and the architecture already decided. Agentic DevOps can place gates where the agent is actively writing, catching a violation the moment it's introduced.

Bazel's build graph shows what the strongest version of this approach looks like. A Bazel build can only see what it has been explicitly declared as a dependency. Package visibility rules turn module boundaries into an API that other modules must be granted access to, and a circular dependency fails the build deterministically. Those are first-class primitives for expressing and enforcing a policy like "X must not depend on Y," and they work.

Where this approach runs out is at the edge of what a rule-writer can anticipate. A pre-commit hook that checks file placement will catch the VS Code utility example only if someone already wrote a rule that anticipated a utility function landing in exactly that wrong location. Rules catch what they were built to catch. Everything else passes through clean.

The rules themselves need governance, and policy encoded as pre-commit or pre-merge checks has to be versioned in Git, reviewed the way code is reviewed, and deployed the way code is deployed. Otherwise the rules drift along with everything else, and a team loses the ability to say what constraints were actually active at the moment a given violation occurred. Enforcement that isn't itself tracked is not enforcement a team can audit later.

Requirements for reviewing structural change

Understanding what an agent built across forty files requires a representation of the codebase that shows how the pieces relate to each other: modules, call graphs, dependency edges, layer boundaries. Without that representation, a change can only be evaluated as a sequence of additions and deletions, which is precisely the limitation this entire argument has been pointing toward.

The shift this requires is from reading what changed to seeing what the system looks like now and checking that against what it was supposed to look like. A living architecture graph, one that maps every module, file, class, function, and call relationship and updates continuously as an agent writes code, lets a developer answer the questions that actually matter in seconds rather than hours: does the new utility function land in the foundation layer, or did it end up next to its caller in violation of the rule? Did the agent introduce a dependency that reaches across a boundary the architecture declared off-limits?

That graph only works as a verification tool if it lives in the repository and versions alongside the code itself. A diagram drawn once in a planning meeting and maintained separately from the codebase is documentation, and documentation diverges from reality the moment an agent makes a sweeping change across dozens of files in a single session. A graph that isn't regenerated with every change isn't describing the system that currently exists.

Constraints expressed on that graph and enforced in continuous integration close the gap that pre-commit rules leave open. A check like gr check, which exits non-zero when the live graph diverges from the declared design, reasons about the structure of the system as a whole rather than checking for one specific, previously anticipated violation at a time. That is the categorical difference from a pre-commit hook: it does not need to know in advance what shape a violation will take, only what shape the architecture is supposed to have.

Understanding Dominant Themes in Reviewing Agentic AI-authored Code identifies the review gaps in agent pull requests as clustering around the structural concerns that inline diff review fails to surface. A graph-based approach to verification targets that category, the one reviewers are currently missing, because the diff never shows it to them even though they are not careless. The role of the developer shifts accordingly, from scrolling through thousands of lines hoping nothing structural slipped past, to steering and verifying against a shared architectural model, the supervisory engineering work that recent research on professional developers describes as the function emerging to replace line-by-line review.

Sources

  1. Understanding Dominant Themes in Reviewing Agentic AI-authored Code
  2. Human-AI Synergy in Agentic Code Review
  3. From Industry Claims to Empirical Reality: An Empirical Study of Code Review Agents in Pull Requests
  4. AIDev: Studying AI Coding Agents on GitHub

More in Agent-Generated Code