Inspiration

After the code is written, a maintainer's day still goes to red pipelines, reviews and deploys. I wanted agents to do that work on a real project — and to see honestly where they fail.

Before → after (Path B)

The original project (BountyPilot, MIT, on GitHub) had no CI at all: tests run by hand, vercel --prod from a laptop, no security scanning, no review, no monitoring. On GitLab it now has a pipeline on every push and MR, three scanners, an agent reviewer, an agent that fixes failing builds, an agent that turns issues into merge requests, review deploys, production behind one human click, releases and a scheduled health check.

Who it's for: solo maintainers and small open-source teams without a release engineer. The agents do the post-code work; the one human keeps the decisions (merge, production).

What it does

  • Pipeline Medic (pipeline failed): reads the failing jobs, quotes the logs, opens the smallest fix as an MR; infrastructure failures become issues; a failed production health check becomes an incident issue. It first checks whether the failure is stale.
  • MR Steward (MR created / marked ready, or mention): reviews the diff, lints CI changes, reads the MR pipeline and security findings, sets one risk::low/medium/high label and posts a "Steward summary" with what a human must check before approving and before deploying. It never approves, merges or deploys.
  • Issue to MR (assign the issue to its service account): posts a plan on the issue (or one clarifying question), implements it with tests on a branch and opens the MR.
  • Pipeline: verify → secure (SAST, secret detection, dependency scanning) → package → review deploy + MCP smoke test per MR → manual production deploy from a version tag → release from the tag notes → scheduled production health check. All nine lifecycle stages have run (table in the README).

How I built it

  • Three custom flows on GitLab Duo Agent Platform (.gitlab/duo/flows/*.yml, flow schema v1), triggered by pipeline, merge-request and assign events.
  • .gitlab-ci.yml with GitLab's SAST, Secret Detection and Dependency Scanning templates, JUnit test reports, a review environment per MR, a manual production gate, releases from tags and a pipeline schedule.
  • The app: BountyPilot, a Node.js MCP server (Streamable HTTP), deployed to Vercel from CI.

Challenges (all visible in the project)

  • The Medic opened four fixes for failures I had already fixed (MRs !3–!6), two of them wrong → it now checks for stale failures, and jobs are interruptible.
  • The Medic made the tests green by spreading a breaking rename of a public MCP tool (MR !10). The Steward first rated it low risk; after I defined the public interface in its prompt, its re-review admitted the first review was wrong and set risk::high.
  • The Builder's first test imported a module that starts a server, so the job hung for an hour (MR !8) → the prompt now warns about import side effects and test jobs time out at 10 minutes; the second try passed (MR !9).
  • Flows do not fire on events another agent created, so the hand-off between agents goes through a human mention.

Accomplishments

  • The Steward found two real defects in a one-line CI change (MR !7) and a breaking API change hidden behind green tests (MR !10).
  • v0.4.0 shipped through the whole lifecycle: tag → manual production deploy → smoke test → release → scheduled health check.

What I learned

  • Green tests are not a correct product; a reviewer agent has to be told what the public interface is.
  • Agents need guardrails against stale work, and a human at the hand-off points.

What's next

Deprecation aliases suggested automatically for breaking renames, a release-notes flow, and a Google Cloud deployment.

Built With

Share this project:

Updates

Submission history