In one minute

Flakeproof is a team of Strands agents that fixes flaky Java tests, with one rule: it only opens a pull request for a fix it can prove.

On a real flaky test in the open-source marine-api library, the gate judged three patches and reran each one 200 times:

Patch Reruns passed Verdict
Put back the "missing" parser (restores the wrong class) 121 / 200 Refused: not proven
@Ignore the test that causes the failure 200 / 200 Refused: band-aid, line 220
The maintainers' real fix 200 / 200 Verified, pull request opened

The @Ignore passed every rerun, and Flakeproof refused it anyway.

Inspiration

You have probably lived this one. A test fails in CI, you rerun the pipeline, and it passes. After the third time, someone adds a retry or an @Ignore and moves on, and the bug is still in the code.

Big teams deal with this too. On September 14, 2026, a GitHub search found 504 issues and pull requests with "flaky" in the title across AWS's own aws organization, 44 of them opened in the last three months. One example is a flaky integration test fix in the AWS SDK for Java from last week. The strands-agents organization has 18 of its own. Atlassian reported that flaky tests cause about 15% of Jira backend build failures and waste over 150,000 hours of developer time a year (Atlassian Engineering).

Tools that repair flaky tests already exist. The hard part is trust: FlakyGuard repairs 47.6% of reproducible flaky tests, and developers accepted 51.8% of those fixes (Li et al.). Coding agents make this harder. An agent told to get CI green can do it by skipping the test, and GitHub's guide to reviewing agent pull requests lists this "CI gaming" as its first red flag (GitHub blog).

I wanted an agent for whoever is stuck on flaky-test duty. It does the slow investigation in the background, and it only comes back with a fix that has evidence attached, or an honest explanation of why it couldn't prove one.

What it does

Flakeproof is for teams running Java test suites in CI. It works on order-dependent flaky tests, where one test leaves shared state behind and a later test fails because of it.

  1. It reruns the flaky test under different test orders to measure how flaky it is before touching anything.
  2. A Strands Swarm of diagnosis agents reads the failure and the source, then runs real experiments: run this test first, then the flaky one, and see if it breaks.
  3. A synthesizer agent records the root cause, and a repair agent edits the code, compiles it and proposes a patch.
  4. A deterministic gate judges every patch twice. It reruns the patch 200 times under rotating orders, and it scans the diff for eight kinds of band-aid, including sleeps, retries, @Ignore and longer timeouts. A patch that passes every rerun is still refused if it is a band-aid.
  5. Only a verified fix reaches the pull request agent, and the pull request carries the evidence: runs before and after, split by test order. Anything else gets a refusal that names the check that failed and the offending line.

The person only steps in at the end, to review a pull request that already has its proof. The web dashboard shows what needs you, every run, and the full evidence behind each verdict.

How I built it

Flakeproof is one Strands Agents Graph: intake, a diagnosis Swarm, a synthesizer, a repair agent, the gate, and then either the pull request agent or a refusal. Every agent runs on Amazon Nova 2 Lite on Amazon Bedrock.

The gate is a custom MultiAgentBase node inside the graph, so the agents cannot route around it. The only edge into the pull request node fires when the gate has verified a candidate. A RefusalGuard hook adds a second lock: when the pull request tool is called, it re-reads the verdict from the database and cancels the call unless the candidate is VERIFIED. Both locks have unit tests that run without a model.

The diagnosis Swarm has triage, test-order, async and resource specialists with handoff limits and timeouts. The agents share 12 @tool functions, including run_pair and run_victim_alone, which launch real JVM runs. TraceHooks write every tool call to SQLite. The band-aid judge uses structured output. It can add a refusal but can never overrule the deterministic scanner, and it never sees the agents' diagnosis.

The rerun harness reads the Surefire XML reports instead of the Maven exit code, deletes stale reports before every run, and counts a skipped test as a failure. Pull requests go through the GitHub REST API with the patched file uploaded byte for byte. The web UI is FastAPI and React, designed by my teammate Basudev Biju, reading the same SQLite record. The repo has 49 unit tests that run in GitHub Actions, and the live dashboard runs on Amazon EC2 at https://flakeproof.me/dashboard.

Challenges I ran into

My first full agent run got it wrong. Triage ran the right test class first, but that class resets its state before each test, so the whole class passed and triage dropped the real suspect. It went on to use 88 pairings on other classes. The synthesizer settled on the wrong root cause, and the repair agent wrote a fix that the model judge rated 95% likely to be real. The gate refused it at 4 of 5 reruns. I changed run_pair to pin a class's methods one at a time when the class as a whole passes, gave the team a budget of 12 pairings, and took real class names out of the tool descriptions. On the next run, triage found the polluting method with a single run_pair call. When I later re-judged the agents' fix on its own, it passed 200 of 200 reruns.

The model judge changed its mind. It called the same wrong patch a 95% real fix in one run, and a 90% band-aid in the next, after reading the agents' wrong diagnosis. That is why a model in Flakeproof can only add refusals, and why the judge no longer reads the diagnosis.

My own statistics were wrong at first. I wrote that 200 clean reruns put the failure rate below 1.5%. Then I split the runs by test order. In alphabetical and filesystem order the flaky test runs before the test that breaks it, so half the runs could never have failed. The confidence bound now counts only the orders where the unfixed code failed every time: 50 of 50 runs, below 6%. It is a weaker number, and a true one.

Windows fought me the whole way. The first pull request rewrote the entire file (+310/-304) because Python turned CRLF line endings into LF, so now the file is uploaded exactly as git stores it. PowerShell's > wrote a patch file as UTF-16 and broke the loader, and the target project will not build on JDK 25.

Accomplishments that I'm proud of

  • A patch that passed 200 of 200 reruns, and the gate still said no.
  • A real pull request whose evidence anyone can check: one commit, six lines, the same fix the maintainers merged upstream in marine-api PR #109.
  • The agents found the cause with one experiment, and their fix held for 200 reruns.
  • A README that says plainly what is verified and what is not.

What I learned

A green build can hide a bug, so the verdict has to look at what the patch changed, not only whether tests pass. Models are useful for diagnosis and for a second opinion, but the final call belongs to deterministic code. I also learned to split my own numbers by condition before believing them.

What's next

  • More order-dependent flaky tests from IDoFT, the International Dataset of Flaky Tests.
  • A CI trigger, so Flakeproof starts on its own when a test flakes.
  • A different rerun strategy for async and timing flakes.
  • Deploying the gate on Amazon Bedrock AgentCore.

Verified, and not yet

Verified:

  • Three patches, three different verdicts, 200 reruns each.
  • A real pull request.
  • The refusal path, with no pull request.
  • The agents finding the cause and writing a fix that passed 200 of 200.
  • The web UI running on the recorded data.

Not yet:

  • A handoff between swarm specialists. Triage solved the case alone every time.
  • More than one target project.
  • AgentCore.

Built With

Share this project:

Updates

Submission history