Customers describe symptoms. Patch ships the fix.

Inspiration

Every software has bugs, and users somehow always manage to find them. In a 2017 QualiTest survey of more than 1,000 US app users, 78% said they notice bugs in the apps they use, and 29% said they notice them at least once a week. The cost of that is direct: 88% said bugs could drive them to abandon an app, and 51% would quit an app entirely if they hit one or more bugs a day. Across the US economy, CISQ puts the annual cost of poor software quality at $2.41 trillion.

On the other side of every bug is a developer, and there are roughly 36.5 million of them worldwide (SlashData). Research by Undo and Cambridge Judge Business School found that developers spend 26% of their time reproducing and fixing failing tests. That is about 620 million hours and $61 billion in salary every year. An average software failure takes 13 hours to fix. Most importantly, 41% of developers named reproducing the bug as the single biggest barrier to fixing it faster, ahead of writing tests and ahead of writing the fix itself.

That last number is the whole thesis behind Patch. The industry has invested enormously in the two ends of the bug lifecycle. Monitoring tools detect errors, and a new generation of AI coding agents writes fixes quickly and cheaply. The middle has barely been touched. Customers report symptoms, not causes. They describe what they saw, often blame the wrong thing, and can't tell you which conditions actually matter, because they never saw the root of the problem. Many real bugs throw no error at all, so monitoring stays silent. An AI coding agent handed a vague complaint will confidently fix the wrong thing, because a complaint isn't a specification.

Reproduction is the bottleneck. Patch automates it fully, creating a QA team for everyone.

What it does

The simple version. Sometimes an app breaks. The person using it says "it broke!" but can't explain how. Patch is a helper that figures out how to break it again the exact same way. Once you can break it on purpose, you can see what's wrong, fix it, and check that it's really fixed.

The technical version. Patch is a multi-agent system that converts unstructured customer complaints into reproduced, evidence-backed defects, then into verified code fixes. It works in five stages:

  • A customer agent extracts testable claims from tickets and screenshots.
  • A reproduction swarm recreates the reported behaviour in controlled real browsers, with network faults injected where needed, and produces a failing regression test.
  • An implementation agent patches the code against that test in an isolated worktree.
  • An independent verifier runs the test against both the original code and the fix branch.
  • A human approves the change before it goes anywhere.

Every claim, experiment and decision is recorded in an append-only evidence ledger. Each role runs on the model family best suited to it, and the system enforces that the agents who build a fix are never the same model family as the agents who check it.

What it looks like from a developer's seat. You connect Patch to your codebase and your support queue. From then on, when complaints come in, you don't read them one by one and try to guess what happened. Patch's intake agent turns each report into a structured brief. The brief separates what the customer observed from what they assumed, groups reports that describe the same underlying fault, and keeps apart reports that only look similar.

The reproduction swarm then investigates on its own, and you can watch it live in the Patch console. You see the hypotheses it is testing, the browser experiments it runs, and the log signals it connects to each failure. When it succeeds, you receive a reproduction packet: a failing test that demonstrates the bug, recorded browser sessions, a timeline of the events that led to the failure, and a root-cause analysis pointing at the responsible code. When it can't reproduce a report, you get a specific list of the information still missing, which goes back to the customer, rather than a ticket closed as "cannot reproduce."

From there, the implementation agent writes the fix and the verifier proves it: the test must fail on your current code and pass on the fix branch. What reaches you is a diff with the evidence attached, ready for your review and approval. Your job moves from investigating to reviewing.

How we built it

We treated the swarm as the product, and designed the architecture to keep it honest.

Separation of duties by model family. The customer agent runs on Gemini, chosen for speed and cost at high volume. The execution and implementation agents run on Claude Code. The supervisor and verifier run on different model families from the agents they check. If one model both diagnoses a bug and writes its fix, a single set of blind spots decides both. The console flags any configuration that would collapse that separation.

Rules enforced in code, not prompts. The server rejects any agent message that doesn't cite evidence. Only the verifier can record a verdict. Stage transitions are gated by functions, not instructions. In our own runs, the pipeline correctly refused to advance because a hypothesis had no stated condition under which it would be proven wrong. The browser tool refuses to run an experiment unless the agent first states what it expects to observe, and the tool, not the model, writes the result to the ledger.

Parallel probes. When the supervisor is uncertain which kind of experiment will surface a fault, the swarm forks two or three probes that pursue different strategies at the same time, each on isolated state. A probe that fails is still useful, because it rules out a hypothesis for the whole team.

[start] swarm                    forking 3 probes
[ok   ] probe:repeat_action      Two identical bookings produced 2 reservations
[ok   ] probe:state_after_action Cancelled R-105; still listed afterwards
[ok   ] probe:capacity           4 tables; 6 bookings accepted; availability 4 to 4 left
[ok   ] swarm                    3 of 3 reproduced; fastest was repeat_action

A real browser lab. Playwright runs experiments across Chromium, WebKit and Firefox. A network fault injector can let the server complete a request and then drop the response, the condition behind a large class of "only happens on my phone" reports. Every session is recorded.

Control plane. A Next.js console on Vercel streams the investigation live, backed by an Upstash Redis append-only event log. Each investigation exports as a single audit trail.

A realistic test target. We built Tablewise, a restaurant booking application with eight seeded defects and one deliberate look-alike report that resembles a real bug but isn't one.

Challenges we ran into

Isolating parallel probes was harder than expected. Giving each probe its own user account wasn't enough, because Tablewise (our demo faulty website we tested Patch on) counts availability per time slot, not per account, so probes contaminated each other's results. Isolation has to follow how the data is actually keyed.

We also learned that some of the most damaging bugs are invisible to manual testing. A defect that only appears when a network response is dropped can't be triggered by a person clicking through the site. We redesigned our defect set so that it contains both bugs a user hits in the first minute and bugs only systematic, fault-injected reproduction can find.

Accomplishments that we're proud of

Patch never converts "not reproduced" into "not a bug." It always returns what is missing.

No agent can fake an experiment. Predictions are required before execution, and verdicts are written by the tooling rather than the model.

Our seeded defects punish shallow work. One cancellation bug returns success and updates the record, yet the reservation reappears on refresh. One party-size bug confirms twelve guests and stores eight. Each surface looks correct except the database, and Patch finds them anyway.

In our test harness, parallel probes reproduced defects 2.5× faster than running the same strategies one after another.

What we learned

Reproduction, not repair, is the hard part of fixing software, and it is where automation creates the most value. A rule written in a prompt is only a suggestion, so every constraint that mattered had to move into code. Interfaces can't be trusted as evidence: two confirmation emails are not two reservations. And a system that says "inconclusive" and explains why is more valuable than one that is confidently wrong.

What's next for Patch

Product. In the near term we're extending Patch in four directions:

  • support for any codebase, not just our test target;
  • a larger probe library covering authentication, concurrency, time zones and pagination;
  • memory across incidents, so patterns like "retry bugs present as mobile-only reports" are learned once;
  • direct intake from Zendesk, Intercom and Sentry, so every new report starts an investigation automatically.

Longer term, every reproduced complaint becomes a permanent regression test, so a customer's test suite grows from real-world failures. The same swarm can also run extreme user journeys before release instead of after a complaint.

Business model. We plan to charge a platform fee plus roughly $200 per verified fix. At 13 hours per failure and an assumed loaded engineering cost of $100 per hour, a single bug costs a company about $1,300 today. That is roughly a 6.5× return per fix.

Market. The addressable spend is the $61 billion a year developers put into debugging. We estimate the serviceable segment, customer-facing web and mobile teams, at about $24 billion. We intend to start at a lower pricing and gradually increase to the fair value as we scale up.

Illustrative projection (based on our pricing and adoption assumptions):

  • 2027: 25 paying teams at $24K per year, $0.6M in annual recurring revenue
  • 2028: 150 teams at $36K, $5.4M
  • 2029: 600 teams at $48K, $28.8M
  • 2031: 3,000 teams at $60K, $180M

Billing based on usage would allow us to reach a lowest estimated ARR of $180M, and would mean capturing under 1% of the serviceable market.

Impact. Recovering even 1% of the 620 million hours developers spend debugging each year returns more than 6 million engineering hours to building. It also means fewer customers hearing that their problem couldn't be reproduced.

Built With

  • anthropic
  • browserbase
  • chatgpt
  • claude
  • gemini
  • huawei
  • juiwen
  • openai
  • rox
Share this project:

Updates

Submission history