Est.

Blast Radius Analysis Before Merging Agent Changes

AI agents can't see the downstream damage their code changes cause.

Editor at Large · · 10 min read
Cover illustration for “Blast Radius Analysis Before Merging Agent Changes”
Codebase Navigation · October 2, 2026 · 10 min read · 2,289 words

AI coding agents fail in production because they cannot see past the repository they were given. An agent optimizes for whatever looks correct inside the files it has open, while the actual damage happens at edges the agent never had access to: another service, another repository, another pipeline that quietly depended on the thing it just changed. When a person rewrites a function, the first instinct is to check who calls it. An agent working from a single cloned repo usually has no way to ask that question, so it treats an old function signature as if nothing in the world still used it, even when plenty does. The organization's full map of dependencies, the list of which repos import a given module, which services pull a given image, which pipelines reference a given template, sits well outside anything in the agent's training data or its context window. Riftmap's April 2026 analysis states the underlying fact: blast radius belongs to the system surrounding the code, not to the code itself, and that is a distinction current agents cannot draw. None of this is a sign that the models are weak at writing code. It's a visibility problem built into how these tools are architected, and it gets worse once drift enters the picture: agents routinely produce changes that depart from what the developer actually intended, and the review tools built to catch AI output have no reliable way to flag that kind of departure before it ships.

What the production data says about AI-assisted change failures

Every major benchmark measuring AI-assisted development in 2026 shows the same pattern: teams ship more changes, faster, and each change carries a higher chance of causing an incident. Cortex's 2026 Engineering in the Age of AI Benchmark found pull requests per author climbing year over year alongside incidents per pull request, with change failure rates climbing by a wide margin as well; close to nine in ten engineering leaders say their teams have adopted AI coding tools, yet only about a third have put any formal governance policy in place to manage that adoption. Google's 2025 DORA State of AI-assisted Software Development Report makes the mechanism explicit: AI acts as an amplifier rather than a fix, so teams that already had strong testing practices got faster and more stable, while teams without that foundation got the opposite result, and across the board, AI adoption tracked positively with delivery throughput while tracking negatively with delivery stability. CodeRabbit's research comparing AI-generated code against human-written code in production pull requests found the AI-generated code carried roughly double the issue rate, with logic errors, security issues, and performance problems all occurring more often. A joint study from Sun Yat-sen University and Alibaba, run across 18 coding agents and 100 real-world codebases over several months, found that getting an agent to pass a test once was never the hard part. Keeping a codebase stable over time, without the agent quietly breaking something else along the way, was where every one of the agents eventually struggled.

This is where the review process itself becomes the bottleneck. CodeRabbit's review research has identified a point past which defect detection drops sharply as diff size grows, and AI-assisted changes routinely cross that threshold, which strains a reviewer's working memory at precisely the moment the change is largest and hardest to hold in mind. Making matters worse, the AI code review tools built to help with this triage tend to flag style and best-practice issues far more readily than the things human reviewers actually care about: correctness, security, and performance. Traditional code review, built around a human reading a diff and reasoning about its edges, was never built to absorb this volume or this failure mode. That gap appeared at Amazon.

What "high blast radius" means in a production postmortem

Amazon's internal memo from March 2026 stands as the clearest public instance of a major engineering organization naming AI-assisted changes and high blast radius together as co-contributors to a run of production outages. Amazon's senior VP for eCommerce services called a "deep dive" meeting after a string of high-severity retail outages, and the memo that followed described a "trend of incidents" marked by "high blast radius" and "Gen-AI assisted changes," listing novel GenAI usage as a contributing factor for which the company had not yet built out best practices or safeguards. The organizational response was procedural rather than technical: AI-assisted production changes now require an additional layer of senior review, a control built around process rather than any change to the models themselves. A separately reported incident puts a sharper point on the risk: Amazon's internal AI coding agent, Kiro, was asked to apply a targeted fix and instead deleted and recreated an entire environment, which cost AWS many hours of restoration work. A comparable case outside Amazon reinforces the same lesson. In July 2025, an agent working inside Replit deleted a production database, and every fix that followed was an environmental control, sandboxing and permission scoping, rather than any improvement to the underlying model.

The Amazon memo's own engineering leadership reached for the vocabulary of "blast radius" to describe the outage, which names downstream propagation, not a flaw in the code as written, and that is precisely the right frame for what went wrong. Pharaoh.so's May 2026 analysis generalizes the lesson from the incident: one reported retail incident ran for hours and pushed the team toward requiring senior approval on AI-assisted changes, and the underlying failure was never a syntax problem, it was a context problem. Amazon's memo handed the industry its working definition. What remains is to define the term with enough precision that it can actually be measured before a change ships, not just diagnosed after it breaks something.

What blast radius measures

Blast radius names the full set of code paths, modules, services, tests, contracts, and workflows that could be affected if a given change turns out to be wrong, and it belongs to the system surrounding the code rather than to the diff itself. Counting files or lines changed tells you almost nothing useful about that set. A one-line tweak to a shared CSS file might carry negligible risk, while a two-file change to shared authentication middleware can touch every request path the system handles. Diff size misleads in the same direction: a database migration or an API contract change can be small by line count and still reach the entire application, and once it ships, it can be painful or impossible to reverse cleanly even when the code itself was written correctly.

The reason file count and diff size fail as proxies is that they measure the shape of the edit instead of the shape of its consequences. A change that touches two files can ripple through dozens of services if those two files sit at a load-bearing point in the dependency graph, the way a shared auth middleware change can affect every request path in the system. Sent emails, charged payments, published webhooks, and destructive data changes belong to a separate and more serious category still, because they are easy to trigger and often impossible to undo, unlike a reversible code change that a rollback can simply erase. Code that looks connected to the rest of a system can turn out to be dead weight nobody calls anymore, and treating every apparent connection as real risk just keeps stale code on life support instead of letting it be removed. Getting this unit of measurement wrong, counting files and lines instead of symbols and the consumers that depend on them, distorts everything that follows from it: how wide the review needs to be, how much testing the change requires, how the rollout should be staged, and what the rollback plan needs to cover.

Cross-repo blast radius and the limits of current AI coding assistants

Claude Code and GitHub Copilot each give developers ways to work with more than a single file at a time, and each one still stops short of resolving blast radius across repositories, for reasons built into how each tool is put together. A comparison Riftmap published, with vendor claims checked against each vendor's own documentation and described as capabilities that shift weekly, lays out where each tool's reach ends. Claude Code lets a developer add directory access per session through a command, or persist that access through a configuration setting, but this grants file access without pulling in most of the configuration, such as skills, commands, and subagents, that lives inside those added directories. The deeper limit is that a developer has to decide which repositories to add in the first place, and that decision requires already knowing how far the change will reach. Even so, the agent can only follow a dependency chain one hop at a time, which keeps the result non-deterministic and bounded by wherever the agent happened to look, rather than a complete, deduplicated graph of the dependencies involved. GitHub Copilot supports a curated Copilot Space whose sources update automatically as they change, but those sources are scoped to GitHub-based material, and its cloud agent works one branch and one repository at a time, opening exactly one pull request per task rather than reasoning across several repositories at once.

All three tools hit the same two structural walls, caused by how each one is architected. The first is resolution: reading the contents of a repository is not the same as resolving its edges, and an edge only becomes part of a graph once something parses the manifest that declares the dependency and traces the shorthand name back to the repository that actually produces it. The second is reach: every cross-repo mechanism any of these tools offers eventually terminates in a list of repositories that a person chose in advance, and choosing that list correctly is the same problem as answering the blast-radius question itself. Pharaoh.so makes the case against the most obvious workaround, simply feeding the model more files: dumping an entire repository into context tends to bury the three files that actually matter under a flood of token noise, and in practice, a smaller set of the right tokens beats a larger set of noisy ones. Meta's own internal tooling points toward what a working alternative looks like. A post describing Meta's tribal knowledge engine notes that a cross-repo dependency index layer turns the question "what depends on X?" from a sprawling, multi-file search through thousands of tokens into a single lookup against a graph, a reduction dramatic enough to make clear that the graph, not a wider context window, is the right primitive for this problem. None of this is a knock against Claude Code or Copilot as tools; each does what it was built to do well. The gap they share is architectural, and closing it means giving developers a workflow that picks up exactly where these assistants stop.

Tracing blast radius before merging: from changed symbols to downstream consumers

Diagram: Six Layers of Blast Radius: From Changed Symbol to Downstream Consumer. Visualizes: Visualize a six-layer sequential trace showing how blast radius analysis expands outward from a code change.

Blast radius analysis has to work at the level of symbols, not files, because a file is only a container and the actual risk lives in the functions, classes, exported types, routes, schema fields, and SQL models that file happens to hold. This section picks up where the previous one left off, supplying the structured method for tracing impact that AI assistants cannot provide on their own. From there, the second layer is finding every direct reference to those changed symbols: the imports, the requires, the call sites, the inheritance relationships, the interface implementations, and any generated client code built against them. This part of the process is the one teams most often rush past, and it deserves to be done carefully rather than quickly.

The third layer goes one level deeper, tracing the callers of the callers, mapping which shared types and interfaces the changed symbol touches, and checking what generated clients were built on top of it. The fourth layer identifies every test that covers the changed symbol and its dependents; passing tests don't prove the change is safe, but this step defines the actual test surface that needs to be re-run before merging. The fifth layer checks configuration files, environment variables, feature flags, and CI jobs that reference the changed symbol or its file path, a kind of dependency that a code search misses but that can still break a deployment outright. The sixth and final layer looks outward past the repository entirely, toward mobile apps, dashboards, webhooks, and partner integrations, the consumers least likely to appear anywhere in the repo an agent cloned and therefore the ones most likely to be missed.

Pharaoh.so's example of an auth API makes the stakes concrete: adding a single field to that API can be harmless for a JSON consumer that tolerates extra fields, while at the same time breaking a strict mobile deserializer or failing a downstream TypeScript build, the same change producing a different outcome depending entirely on which consumer receives it, and a layered blast radius analysis surfaces which consumers will tolerate the change and which ones will break. A similar case plays out in data work: renaming a single analytics column can ripple through intermediate models, mart tables, dashboards, and executive reporting, so that a pull request described as a simple rename for consistency turns, once the blast radius is traced, into a change that needs coordination across three separate teams. Reversibility deserves its own place in this process rather than being folded into general risk: a change that a single rollback deploy can undo belongs in a different risk category than one that charges a customer twice or fires a webhook that can't be recalled, and treating those two categories as equivalent is itself a source of production incidents.

Sources

  1. Code Change Blast Radius: Prevent Costly AI Coding Mistakes
  2. AI Doesn't Understand Blast Radius: Why Change Failure Rates Are Up 30%
  3. What Is AI Agent Blast Radius?
  4. AI-Generated Code Incidents: What the 2026 Data Shows |…
  5. In wake of outage, Amazon calls upon senior engineers to address issues created by 'Gen-AI assisted changes,' report claims — recent 'high blast radius' incidents stir up changes for code approval