Inspiration

Code generation is solved — but everything that happens after the code is written still eats a senior engineer's afternoon: the careful security review, the "are there tests for this?" check, the dependency hygiene nobody does. We wanted agents that take over that post-code work without taking over the decision. So we built a supervised crew: deterministic logic decides, the agents narrate and propose, and the human approves the outcome.

What it does

MergeGate is a supervised DevSecOps agent crew that runs on every merge request. A deterministic verdict engine parses SAST (semgrep, bandit), pytest, secret-scan, and dependency-policy results and emits a unified PASS/WARN/FAIL verdict. Sentinel triages every security finding and proposes fixes as suggestion blocks. Scribe maps the diff to test coverage and proposes the missing tests. Steward enforces policy: no hard-coded secrets, dependency pins above the policy floor, license headers, MR hygiene. A FAIL verdict turns the pipeline red; the author accepts suggestions, requests changes, or records an explicit waiver — then merges.

How we built it

Path A (Start Fresh), Supervised autonomy: a small Flask demo app with deliberately seeded, fully disclosed vulnerabilities as the fixture the crew reviews. Real .gitlab-ci.yml: build → test → sast → review → package → deploy-dry-run, with the review stage running only on merge requests. The verdict engine is pure-stdlib Python with fixed severity mappings and policy thresholds — no LLM in the decision path, following the winning pattern from prior GitLab hackathons (deterministic analysis, LLM narration). We built a custom GitLab Duo flow "MergeGate Review" (4 AgentComponents: sentinel → scribe → steward → synthesis) using the flow registry v1 schema, triggered by @-mention via a Mention trigger. Agent prompts live in .gitlab/duo/ with hard boundaries: agents propose via suggestions; they never commit, approve, or merge.

Challenges we ran into

The Duo Agent Platform's custom flow schema had to be extracted verbatim from the flow registry v1 spec — our first attempt used a guessed prompts format and was rejected; the correct schema uses prompt_id, prompt_template (system/user), and params. The Monaco editor staircases multi-line pastes, so we shipped the flow definition as single-line JSON. Keeping the demo honest: every seeded flaw is labeled as a fixture in code, README, and agent prompts, so nothing is ever presented as a real discovery. Making "supervised" demoable: the approval moment — not the automation — is the product, so the video centers on the verdict and the human's accept.

Accomplishments that we're proud of

A genuinely deterministic verdict engine (23 passing tests) that gates a real pipeline — the agents argue from evidence, not vibes. A live custom Duo flow that ran end-to-end: @-mention triggered session 9098473, all three agents independently returned FAIL (Sentinel: 7 security findings including 3 critical; Scribe: 7 untested paths; Steward: 8 policy findings), synthesis posted "FAIL — Do Not Merge" as an MR note, and the human author accepted the verdict and declined the merge. Three tight agent definitions with explicit boundaries, plus repo-wide Duo chat rules.

What I learned

Supervised autonomy is the sweet spot for trust: developers don't want agents that merge — they want agents that do the homework and show their work. Deterministic-first design makes agent output reviewable: when the logic is fixed and visible, the narration becomes useful instead of suspicious. Hackathon honesty compounds: labeling every fixture up front made the judging story stronger, not weaker.

What's next

Wire the MCP servers (GitLab, issue tracker) so the crew can file and link follow-up issues automatically. Extend Steward's policy set (license scanning, container-image policy). Push toward the remaining lifecycle stages: plan (issue-driven flows), release, configure, monitor.

Built With

Share this project:

Updates

Submission history