PatchPilot

Prove it's exploitable. Patch it. Prove it's dead.

PatchPilot is an agent that checks whether a security bug in one of your dependencies can actually be exploited, fixes it, repairs whatever the upgrade breaks, and then re-runs the attack to show the bug is gone.

Inspiration

Security scanners just match version numbers. "You have PyYAML 5.3.1, it's in the vulnerable range, here's a ticket." They can't tell you whether that bug is actually reachable in your app. So teams get a pile of alerts that are mostly noise, and end up patching nothing.

And patching is scary on its own: bumping a package can break things in production. So people put it off. The median time to fix a known vulnerability is about 252 days.

Two problems, and nobody solves both together:

  1. Is this bug actually exploitable here, or just noise?
  2. Who fixes the code the upgrade breaks?

Dependabot opens the PR and leaves the breakage to you. Infield plans the upgrade but a human still writes the fix. None of them prove the bug is real by actually exploiting it. We wanted to do the whole thing, end to end.

What it does

PatchPilot runs a loop that corrects itself:

  1. Detect — pip-audit scans the repo and finds a real CVE.
  2. Prove — in a sandbox, the agent runs an actual exploit. For our demo it sends a malicious YAML config file and shows the server executes the hidden code — remote code execution.
  3. Upgrade — it bumps the package to the patched version.
  4. Observe — it runs the tests and collects whatever broke.
  5. Fix — the model repairs the broken code from those errors.
  6. Verify — it re-runs the tests, looping back to fix if needed (up to 3 tries).
  7. Guard — a check makes sure it didn't weaken security just to pass a test.
  8. Re-prove — it runs the same exploit again; now the malicious file is blocked and nothing executes.
  9. Gate & merge — it opens a PR, an independent AI reviews it, and it only merges if both the tests and the review pass.

If it can't fix something after three tries, it doesn't hide it — it opens a draft PR explaining what it tried and hands it to a human.

How we built it

The core is a Python agent built as a simple state machine, so each step of the loop is easy to follow and safe to re-run.

  • Daytona runs the sandboxes. One sandbox runs the exploit before and after; another runs the fix loop. Anything risky — installing a vulnerable package, running an exploit, running the model's code — happens in there, never on our own machine.
  • Fireworks AI runs two models: a small fast one to read errors, and a coder model to write the fixes. The exploit itself is a real, vetted script the agent runs in the sandbox before and after the patch. It's fast enough to do three fix attempts in under a minute.
  • Braintrust records every step and scores each run, including a check for our biggest worry: did the agent weaken security to make a test pass?
  • CodeRabbit is the independent reviewer on the PR. Nothing merges unless both the tests and CodeRabbit approve. The agent that writes the fix never gets to approve it.
  • GitHub APIs handle the branch, commit, PR, and checks.

Our demo bug

We built the loop around CVE-2020-14343, a PyYAML remote-code-execution bug. PyYAML is the standard tool Python apps use to read YAML config files. In version 5.3.1, reading a document with yaml.load() can be tricked into running code hidden inside it — so a malicious config file can take over the server.

It's a great example because the fix forces a jump to a new major version (6.0), and that jump breaks the code in a real, fixable way — and the break is the security fix:

# Before (PyYAML 5.3.1) — runs code hidden in a malicious YAML document
data = yaml.load(text)

# After (PyYAML >= 6.0) — yaml.load now REQUIRES you to name a loader;
# the secure choice, safe_load, blocks the attack
data = yaml.safe_load(text)

On 6.0, yaml.load(text) errors until you pick a loader — so the agent is forced to make a security decision. The secure choice (safe_load) closes the hole; the lazy choice (yaml.Loader) would pass the tests but keep the RCE alive. So the agent isn't just clearing an error; it's actually reasoning about the security fix.

Challenges we ran into

  • Keeping the exploit real. We could have faked the "accepted / blocked" result. Instead we use a real scanner to detect the bug and run the same exploit before and after, so the result is earned.
  • The tempting shortcut. When a test fails, the easy way to "fix" it is to use the unsafe loader (yaml.Loader) instead of the safe one — it passes the tests but keeps the RCE alive. Models will do exactly that. Catching it became one of our favorite features.
  • Stopping the loop. Fix loops can run forever. A hard limit of three tries plus a handoff to a human gives it brakes.
  • Getting the framing right. The hard part wasn't the code — it was not building "just another bug-fixing agent." The point is proving a bug is real and proving it's dead, and every choice had to serve that.

What we learned

  • Running the exploit beats reasoning about it. Analyzing a patch is an opinion. Running the attack before and after is proof.
  • Splitting the roles matters. Making the agent that writes the fix unable to approve it kills the "the AI graded its own work" problem.
  • The scary step is the valuable one. Bumping a version is easy; fixing what it breaks is the reason people avoid patching — so that's the part worth automating.

What's next

  • More bug types beyond config-injection / RCE, each with its own exploit.
  • Handling upgrades that force other packages to upgrade too.
  • Support for more languages like JavaScript and Go.
  • Triggering from real scanner alerts and CI failures instead of a demo repo.
  • A benchmark across many repos, so our success rate is measured, not guessed.

Built With

  • agents
  • ai-agents
  • braintrust
  • coderabbit
  • daytona
  • fastapi
  • fireworks-ai
  • github-api
  • langgraph
  • llm
  • pip-audit
  • pyjwt
  • pytest
  • python
  • security
Share this project:

Updates

Submission history