Inspiration

Every production change carries consequences that aren't visible in the diff.

You change a transaction fee in one service, but a refund path, a legacy billing component, or a configuration default may each carry their own copy of that number. Engineers spend hours tracing these consequences across source, tests, documentation and dependencies, yet some still get missed, and the ones that do can become incidents.

Existing tools can tell you what might be affected. I wanted to explore a different question: what happens if you stop asking an AI to predict the blast radius and instead make agents investigate the change, construct an experiment, and then let the system execute it?

A software change is a hypothesis about behaviour. Replay tests it.

What it does

Replay takes a proposed change in plain language, such as:

"Change the transaction fee from 2.5% to 2%."

It then rehearses that change before it ships:

  1. Maps where the change can reach, including places that encode the same rule in a different form
  2. Predicts what each relevant behaviour should do, grounding every prediction in a written rule
  3. Executes those scenarios against the current system
  4. Applies the proposed change inside an isolated copy of the repository
  5. Executes the same scenarios again
  6. Compares what actually changed, rather than what the model expected
  7. Challenges the findings with an adversarial verifier whose job is to prove the conclusion wrong

Because each scenario commits to a prediction before execution, the result falls naturally into four cases:

Observed: unchanged Observed: changed
Predicted: changed Under-propagation — the rule never reached here Confirmed effect
Predicted: unchanged Contained Over-propagation — the change leaked

On the demo application, one configuration change produces both failure modes. The refund fee moves even though a contract fixes it at 2.5%, while the legacy billing engine stays at its old value even though the same business rule says it should move.

Neither behaviour appears in the diff, and neither raises an error.

Replay never modifies anything consequential. It produces an evidence-backed risk report and stops for human approval.

The experiment is AcmePay

The fee change isn't hypothetical. It runs against AcmePay, a small payments application I built alongside Replay specifically to exercise these failure modes.

AcmePay includes a payment service, refund service, modern billing service and a legacy billing engine ported from an older system, all working against a shared set of business rules.

The code deliberately contains realistic inconsistencies of the kind Replay is designed to find, and it ships in the same repository so anyone can inspect the implementation, run the experiment and trace the numbers back to the source:

https://github.com/bzdmin/replay/tree/main/demo-repo/acmepay

How I built it

Replay is built with the Strands Agents SDK and runs its agents on Amazon Bedrock. There are three agents, each with a specific responsibility:

Agent Responsibility
Impact Agent Traces what the change reaches and where it may not propagate
Scenario Agent Designs falsifiable predictions about behaviour
Verifier Agent Adversarially challenges the findings

They use real repository tools including search_repository, read_file, list_repository, search_tests and read_documentation.

The Replay Runner is deliberately not an agent.

Executing a scenario and comparing two outputs are deterministic operations, so putting a model in that path would add cost, latency and another source of uncertainty. The numbers in the report should come from something that actually ran.

The sandbox keeps the experiment isolated by running scenarios as real subprocesses inside two copies of the repository, one untouched and one with the proposed change applied. Replay captures the actual return values and exit codes, so when it claims that behaviour changed, there is an execution trace behind the claim.

The application runs as a container and streams progress over server-sent events, letting you watch the rehearsal happen instead of staring at a spinner. A rehearsal costs about $0.58 in model tokens, measured from the agents' own usage metrics.

Challenges I ran into

The agents answered from imagination. My first implementation asked Strands for structured output directly, which meant the tool-use loop never ran. The agent confidently produced scenarios referencing services.transaction_service, a module that didn't exist. The fix was to separate investigation from serialization: first investigate with tools, then serialize only what was actually found.

Bedrock's error codes sent me in the wrong direction. A model that required the Anthropic use-case form returned ResourceNotFoundException, while an account without access returned AccessDeniedException. In both cases, the message contained the useful explanation, not the exception name. I eventually wrote a preflight check that surfaces the actual message before the agents run.

The Impact Agent was too helpful. It kept fixing the legacy billing engine on its own initiative, which sounded sensible but destroyed the experiment. If the agent repairs the consequence of the engineer's change, there is nothing left for Replay to discover, so I constrained it to reproduce the proposed change faithfully, including its shortcomings.

The Scenario Agent predicted from the code instead of the rules. It reasoned that the legacy engine had its own constant, predicted that it would stay unchanged, and was technically correct. But that meant the system had predicted the bug instead of testing the requirement. It now has to ground predictions in named business rules, because "it has its own constant" is not a rule.

The progress stream died silently in production. The verifier could spend minutes reading code without emitting anything, and the hosting proxy would eventually reset the connection because nothing had travelled over the wire. The work completed successfully every time, but the browser never received the answer, so I added progress events throughout the long-running stages.

Replay found a flaw in Replay. When I pointed it at its own source, the verifier read the repository state from before the edit and concluded that the change had never been applied. The edit had happened correctly, but the evidence available to the verifier was stale. That exposed an important boundary in the design: verification needs access to the same state that produced the observed result.

What I'm proud of

The evidence is real. Every number on screen comes from executing code twice, inside two isolated copies. Nothing is inferred from a screenshot or generated as an example.

It finds both directions of failure. Most change analysis looks for things that break when they change. Replay also looks for things that fail to change when they should, which can be just as dangerous and is much harder to see in a diff.

It refuses to guess. When I pointed Replay at a codebase outside the patterns it was designed for, many scenarios could not be executed. It discarded them with reasons instead of inventing findings.

The verifier caught a real ambiguity. An edit targeting 2.5% could match multiple locations in the same file, so the verifier demanded manual confirmation. The sandbox now rejects edits whose target text is ambiguous rather than choosing a location on the engineer's behalf.

What I learned

Structure is not evidence. A model can produce a perfectly valid structured response that is completely unsupported. The important distinction is between asking a model for an answer and giving it a process that requires the answer to be grounded in something observable.

That is why Replay separates investigation, prediction, execution and verification. The agents reason about the experiment, but the experiment itself produces the evidence.

The oracle problem changes when the baseline can execute. Instead of asking a model to decide what the new system should return, Replay uses the current implementation as the behavioural baseline and asks the model a narrower question: should this particular observation move according to the written rule?

The interesting result is then the disagreement between the rule-backed prediction and what the proposed version actually does.

Coverage is not the same as useful testing. Research on targeted mutation testing at scale found that concern-targeted generation killed substantially more seeded faults than coverage-driven generation, while some useful tests added no line coverage at all. That changed how I designed the Scenario Agent: its job is to exercise the places where a rule is encoded, not simply to touch more lines.

Measure before you optimise. I was convinced the deployment was running out of memory and nearly moved hosts over it. Measurement showed 104 MB against a 512 MB limit. The real problem was a long-lived network connection.

Name your limits. Replay needs behaviour it can execute, intent written down, and behaviour reachable in a single call. Stateful workflows exposed that third requirement clearly. Saying what the system cannot currently prove is more useful than pretending it can.

What's next for Replay

  • Scenarios with setup, so stateful behaviour can become reachable without sacrificing sandbox isolation
  • Beyond code, including configuration, dependency upgrades and schema migrations
  • Behavioural equivalence for migrations, where the question is whether a rewritten service preserves the behaviour that mattered
  • Pull request integration, so a rehearsal can happen before review rather than after a regression reaches production

Built With

Share this project:

Updates

Submission history