Inspiration
We've all experienced a commit or PR which looks fine on the surface, all test pass, it seems to work properly, but it causes regressions in unexpected places. Judging the full blast radius of a change and performing thorough QA testing is a significant cost to companies, and one which has thus far resisted automation.
The build was green. The checkout was lying.
A customer enters a discount code. The website says “Coupon applied.” The request succeeds. Nothing crashes.
But the price never changes.
We’ve all seen some version of this: a pull request looks reasonable, passes its checks, and quietly breaks something a developer didn’t think to test. Understanding the full impact of a change still means opening browsers, repeating journeys, comparing results, and tracing failures back to code.
We built Aftershock to take on that investigation.
Our question was simple: What if every code change came with a team that checked what you meant to build—and what you accidentally broke?
What we built
Aftershock is a six-agent QA and repair team that turns code changes into browser-tested evidence. It reads a change, tests deployed user journeys, compares baseline and preview behavior, reproduces suspected failures, and attempts a source-code repair.
Meet the team: our six-agent Avengers:
- Diffany - The Scout. Reads the commit and diff, identifies relevant routes, and turns the developer’s intent into testable assertions.
- QAizen - The Feature Tester. Operates real browsers to check whether the feature actually delivers its promise.
- Doppler - The Regression Tester. Runs the same journey against baseline and preview deployments to expose unintended changes.
- Gavel - The Critic. Challenges the initial reports, uses fresh-session reproductions, and decides which findings have enough support to file.
- Clueso - The Investigator. Connects the failure and its evidence to likely causes in the source.
- Patchouli - The Repair Agent. Uses Codex to edit an isolated checkout. The pipeline then deploys the patch and tests it before publishing the repair result.
Browserbase provides the real browser sessions. Stagehand powers the interactions. Our dashboard makes the investigation visible.
You can follow the assignments, inspect captured screenshots, compare the two deployments, read the patch, and see whether verification actually passed.
It actually works
We built Meridian, a small storefront with deliberately planted bugs, as a controlled test subject.
The bugs are known. The browser interactions, deployments, generated patch, and verification are real.
The feature commit added coupon codes at checkout. A customer adds an $84 scarf, enters SAVE20, and sees a success message.
The expected total is:
$$ \$84.00 \times (1 - 0.20) = \$67.20 $$
The actual total stays $84.00.
That is the feature failing its own promise.
But the more interesting failure is somewhere else.
The same commit changed a shared price-formatting helper. The cart—outside the stated coupon feature—now displays $NaN where the baseline displays $84.00.
We added coupons. We accidentally broke the cart.
Aftershock caught both.
| Stage | What happened in the verified run |
|---|---|
| Diffany | Planned five assignments |
| QAizen | Caught the coupon-total failure |
| Doppler | Caught the cart regression against the baseline |
| Gavel | Confirmed both findings through reproduction |
| Issue filing | Created GitHub issues #22 and #23 |
| Clueso | Identified lib/money.ts as the likely cause of the cart failure |
| Patchouli | Generated the cart repair on its first attempt |
| Curtain Call | Replayed the failed journey against the repaired deployment; verification passed |
| Publication | Opened verified repair PR #24 |
Three minutes. Two confirmed bugs. One source-code repair tested on a real deployment.
The fix was small:
-export function formatPrice(cents: number, discountPct?: number): string {
- const net = cents * (1 - discountPct! / 100);
+export function formatPrice(cents: number, discountPct: number = 0): string {
+ const net = cents * (1 - discountPct / 100);
return `$${(net / 100).toFixed(2)}`;
}
When a caller omitted the optional discount, the calculation used undefined and produced NaN. Defaulting the discount to zero restored the original behavior.
The impressive part wasn’t generating two lines of code. It was closing the loop:
Observe the failure → reproduce it → change the source → deploy the change → replay the failed journey → verify the result.
The separate coupon-total bug remained open. Aftershock currently repairs the highest-ranked finding per run. We distinguish a verified repair from a claim that the entire application is bug-free.
Original feature PR · Verified repair PR
How we built it
Aftershock combines TypeScript, Next.js, Browserbase, Stagehand, OpenAI models, the Codex SDK, GitHub, and Vercel preview deployments.
At its core are two complementary ways to test a change.
Challenges we ran into
A page can differ between loads even when its code hasn’t changed. Timestamps move, session IDs change, and accessibility nodes get renumbered. That makes “did this actually change?” a harder question than it sounds. Get it wrong, and the useful findings disappear into noise.
So we built a canary: run the same journey against the same deployment and require it to produce no meaningful differences. It exposed session-specific node IDs in Stagehand’s accessibility snapshots, identical content could look different simply because its elements had different identifiers. We used these clean-versus-clean runs to validate and refine our comparator.
Filtering out noise without filtering out the bug took more care than we expected. Ignoring numbers would make comparisons quieter, but $84.00 versus $67.20 is exactly the evidence behind our coupon failure. We deliberately avoided blanket numeric filtering, targeting known sources of variation such as timestamps, session identifiers, build hashes, and cache-busting parameters instead.
Even after that, an empty iframe in one capture caused verification to reject a working cart repair. The price was fixed; the comparator was still reporting an irrelevant difference. That taught us that the system checking the fix needs just as much scrutiny as the agent writing it.
Navigation was another surprisingly important detail. An instruction like “go to /checkout” should navigate directly, rather than depend on finding a matching link on the current page. Making explicit navigation deterministic removed an unnecessary source of variation.
The first complete integration was harder than building any individual agent. We had to connect Codex, Browserbase sessions, isolated checkouts, Vercel previews, and the GitHub API using real credentials, and make failures visible at every handoff.
A generated patch, a deployed patch, and a verified repair are three different milestones. Getting all three to happen in sequence was the real challenge.
Accomplishments that we're proud of
The loop closes on a real repository. A commit went in through, and a verified pull request came out the other end, with three issues filed along the way. Nobody made a decision in between and nobody opened a browser.
The retry fired on its own. The first patch failed verification, that failure was handed back to Codex as context, and the second attempt passed. We spent more time on that path than on the happy one, because a system that can't tell when its own fix didn't work isn't much use.
Codex is part of the product rather than just how we built it. It runs inside the pipeline on a real checkout of your repo, and the constraints it's given are checked afterwards rather than trusted.
What we learned
We wrote three of the agents and about ninety tests in first few hours, and it turned out not to matter very much. None of it had run against a real Codex process, a real browser or a real repository. The first time we tried that, most of it broke, and fixing it took the rest of the day. If we started again we'd do the first full run with real credentials in hour one, on whatever half-built thing existed then, and build outwards from there.
We also expected driving the browsers to be the hard part. It wasn't. The hard part was deciding whether what came back meant anything, because two loads of the same page disagree constantly. Every hour we spent on that filter was worth more than an hour spent on any of the agents.
What’s next for Aftershock
- Simpler onboarding: GitHub installation and deployment configuration that take less setup.
- Authenticated journeys: testing applications beyond what anonymous visitors can reach.
- Multiple repairs per run: addressing more than the highest-ranked confirmed finding.
- Persistent regression tests: leaving behind an executable test for each confirmed failure.
- Broader evaluation: measuring reliability across applications beyond our controlled storefront.
- Learning from developer feedback: using accepted findings and dismissed noise to improve future runs.
Every code change has an aftershock. We’re building the team that finds it.
Built With
- browserbase
- codexsdk
- github
- next.js
- openai
- stagehand
- typescript
- vercel
Log in or sign up for Devpost to join the conversation.