-
-
Four ways to rehearse a change: three real presets or your own change, run live
-
Six stages of a real rehearsal, from investigation to verification, with live execution evidence
-
The diff looked safe. The behaviour wasn't. Every value came from executing both versions, not reading the diff
-
A separate verifier challenges the findings, confirming 6 of 6 divergences against the evidence
-
Three agents, one deterministic runner. Reasoning is agentic; execution and comparison stay deterministic.
-
AcmePay is the real, runnable application behind the rehearsal, with its services and business rules exposed
Inspiration
Every production change carries consequences that aren't visible in the diff.
You change a transaction fee in one service, but a refund path, a legacy billing component, or a configuration default may each carry their own copy of that number. Engineers spend hours tracing these consequences across source, tests, documentation and dependencies, yet some still get missed, and the ones that do can become incidents.
Existing tools can tell you what might be affected. I wanted to explore a different question: what happens if you stop asking an AI to predict the blast radius and instead make agents investigate the change, construct an experiment, and then let the system execute it?
A software change is a hypothesis about behaviour. Replay tests it.
What it does
Replay takes a proposed change in plain language, such as:
"Change the transaction fee from 2.5% to 2%."
It then rehearses that change before it ships:
- Maps where the change can reach, including places that encode the same rule in a different form
- Predicts what each relevant behaviour should do, grounding every prediction in a written rule
- Executes those scenarios against the current system
- Applies the proposed change inside an isolated copy of the repository
- Executes the same scenarios again
- Compares what actually changed, rather than what the model expected
- Challenges the findings with an adversarial verifier whose job is to prove the conclusion wrong
Because each scenario commits to a prediction before execution, the result falls naturally into four cases:
| Observed: unchanged | Observed: changed | |
|---|---|---|
| Predicted: changed | Under-propagation — the rule never reached here | Confirmed effect |
| Predicted: unchanged | Contained | Over-propagation — the change leaked |
On the demo application, one configuration change produces both failure modes. The refund fee moves even though a contract fixes it at 2.5%, while the legacy billing engine stays at its old value even though the same business rule says it should move.
Neither behaviour appears in the diff, and neither raises an error.
Replay never modifies anything consequential. It produces an evidence-backed risk report and stops for human approval.
The experiment is AcmePay
The fee change isn't hypothetical. It runs against AcmePay, a small payments application I built alongside Replay specifically to exercise these failure modes.
AcmePay includes a payment service, refund service, modern billing service and a legacy billing engine ported from an older system, all working against a shared set of business rules.
The code deliberately contains realistic inconsistencies of the kind Replay is designed to find, and it ships in the same repository so anyone can inspect the implementation, run the experiment and trace the numbers back to the source:
https://github.com/bzdmin/replay/tree/main/demo-repo/acmepay
How I built it
Replay is built with the Strands Agents SDK and runs its agents on Amazon Bedrock. There are three agents, each with a specific responsibility:
| Agent | Responsibility |
|---|---|
| Impact Agent | Traces what the change reaches and where it may not propagate |
| Scenario Agent | Designs falsifiable predictions about behaviour |
| Verifier Agent | Adversarially challenges the findings |
They use real repository tools including search_repository, read_file,
list_repository, search_tests and read_documentation.
The Replay Runner is deliberately not an agent.
Executing a scenario and comparing two outputs are deterministic operations, so putting a model in that path would add cost, latency and another source of uncertainty. The numbers in the report should come from something that actually ran.
The sandbox keeps the experiment isolated by running scenarios as real subprocesses inside two copies of the repository, one untouched and one with the proposed change applied. Replay captures the actual return values and exit codes, so when it claims that behaviour changed, there is an execution trace behind the claim.
The application runs as a container and streams progress over server-sent events, letting you watch the rehearsal happen instead of staring at a spinner. A rehearsal costs about $0.58 in model tokens, measured from the agents' own usage metrics.
Challenges I ran into
The agents answered from imagination. My first implementation asked Strands
for structured output directly, which meant the tool-use loop never ran. The
agent confidently produced scenarios referencing
services.transaction_service, a module that didn't exist. The fix was to
separate investigation from serialization: first investigate with tools, then
serialize only what was actually found.
Bedrock's error codes sent me in the wrong direction. A model that required
the Anthropic use-case form returned ResourceNotFoundException, while an
account without access returned AccessDeniedException. In both cases, the
message contained the useful explanation, not the exception name. I eventually
wrote a preflight check that surfaces the actual message before the agents run.
The Impact Agent was too helpful. It kept fixing the legacy billing engine on its own initiative, which sounded sensible but destroyed the experiment. If the agent repairs the consequence of the engineer's change, there is nothing left for Replay to discover, so I constrained it to reproduce the proposed change faithfully, including its shortcomings.
The Scenario Agent predicted from the code instead of the rules. It reasoned that the legacy engine had its own constant, predicted that it would stay unchanged, and was technically correct. But that meant the system had predicted the bug instead of testing the requirement. It now has to ground predictions in named business rules, because "it has its own constant" is not a rule.
The progress stream died silently in production. The verifier could spend minutes reading code without emitting anything, and the hosting proxy would eventually reset the connection because nothing had travelled over the wire. The work completed successfully every time, but the browser never received the answer, so I added progress events throughout the long-running stages.
Replay found a flaw in Replay. When I pointed it at its own source, the verifier read the repository state from before the edit and concluded that the change had never been applied. The edit had happened correctly, but the evidence available to the verifier was stale. That exposed an important boundary in the design: verification needs access to the same state that produced the observed result.
What I'm proud of
The evidence is real. Every number on screen comes from executing code twice, inside two isolated copies. Nothing is inferred from a screenshot or generated as an example.
It finds both directions of failure. Most change analysis looks for things that break when they change. Replay also looks for things that fail to change when they should, which can be just as dangerous and is much harder to see in a diff.
It refuses to guess. When I pointed Replay at a codebase outside the patterns it was designed for, many scenarios could not be executed. It discarded them with reasons instead of inventing findings.
The verifier caught a real ambiguity. An edit targeting 2.5% could match
multiple locations in the same file, so the verifier demanded manual
confirmation. The sandbox now rejects edits whose target text is ambiguous
rather than choosing a location on the engineer's behalf.
What I learned
Structure is not evidence. A model can produce a perfectly valid structured response that is completely unsupported. The important distinction is between asking a model for an answer and giving it a process that requires the answer to be grounded in something observable.
That is why Replay separates investigation, prediction, execution and verification. The agents reason about the experiment, but the experiment itself produces the evidence.
The oracle problem changes when the baseline can execute. Instead of asking a model to decide what the new system should return, Replay uses the current implementation as the behavioural baseline and asks the model a narrower question: should this particular observation move according to the written rule?
The interesting result is then the disagreement between the rule-backed prediction and what the proposed version actually does.
Coverage is not the same as useful testing. Research on targeted mutation testing at scale found that concern-targeted generation killed substantially more seeded faults than coverage-driven generation, while some useful tests added no line coverage at all. That changed how I designed the Scenario Agent: its job is to exercise the places where a rule is encoded, not simply to touch more lines.
Measure before you optimise. I was convinced the deployment was running out of memory and nearly moved hosts over it. Measurement showed 104 MB against a 512 MB limit. The real problem was a long-lived network connection.
Name your limits. Replay needs behaviour it can execute, intent written down, and behaviour reachable in a single call. Stateful workflows exposed that third requirement clearly. Saying what the system cannot currently prove is more useful than pretending it can.
What's next for Replay
- Scenarios with setup, so stateful behaviour can become reachable without sacrificing sandbox isolation
- Beyond code, including configuration, dependency upgrades and schema migrations
- Behavioural equivalence for migrations, where the question is whether a rewritten service preserves the behaviour that mattered
- Pull request integration, so a rehearsal can happen before review rather than after a regression reaches production


Log in or sign up for Devpost to join the conversation.