Inspiration
An issue closes when its merge request merges. The tests that let it merge ran in CI, against CI's configuration. Nobody looks at production, so on most teams "done" means "merged".
The failures users actually hit live in that gap: a server in another time zone, a feature flag nobody switched on, a missing secret, a CDN rule. Now that agents write more of the code, changes merge faster than anyone verifies them, and the gap grows.
We wanted the issue tracker to tell the truth: an issue is done when it works in production, and here is the photo.
What it does
- Proves every deploy. When a pipeline on the default branch passes after a human merge, a GitLab Duo custom flow starts with no human action. Its examiner agent:
- finds the merge requests in the deploy, and the issues they close;
- reads each issue's acceptance criteria;
- looks at the live pages (their roles, names and labels);
- writes a check plan in a fixed vocabulary of 15 step words.
- Photographs production. A deterministic runner, Playwright in Chromium, walks each plan against the live app and photographs every step.
- It first waits until production reports the commit it was asked to prove.
- It can check what a visitor takes away too, for example reading a downloaded calendar file's start time in the right time zone.
- Lets only the runner decide. There are three outcomes:
- held: every step happened as written;
- failed: a step didn't, with the exact sentence of what was seen;
- needs a person: a browser can't observe it, like an email arriving. Never passed silently.
- Files the evidence in GitLab. The clerk agent handles each issue:
- held: closes it with an inline strip of frames;
- failed: reopens it with the failing frame;
- needs a person: labels it for one.
- Fixes what it can, under a policy. The fixer agent opens a small fix merge request on a
proof/fix-<iid>branch, with a test. A guard job merges it only if it stays within the policy:- at most 3 files and 40 lines;
- no changes to the pipeline, flow, runner or checks;
- no high or critical SAST or secret findings;
- a green pipeline. Otherwise it writes why and leaves it for a person.
- Re-checks the fix in production. After the fix deploys, a CI job replays the exact plan that failed, with no model involved. If it holds, the issue closes with the new frames.
- Keeps done proven. A scheduled light table job replays every proven check every night and reopens any issue that regresses.
- Keeps a tamper-evident record. Every roll (one run of the checks) goes to the proof-sheet site on Google Cloud Run. Each roll is hash-chained to the one before it.
- Publishing is accepted only from this project's GitLab jobs, verified by asking GitLab.
- Release notes link each deploy to its proof instead of claiming done.
The demo, for real. Carrel is a study-room booking app for a public library in Pune, and the demo product this repository ships. Over three releases on Oct 7, every feature passed its tests, merged and closed its issue, and four of them didn't work in production:
- Floor plans (#7). The Dockerfile never copied the images, so every plan answered 404.
- 12 minutes after the merge, the flow reopened the issue with the broken frame.
- The fixer opened a one-line fix (
COPY assets ./assets), and the guard merged it within policy. - The re-check held and the issue closed, 31 minutes after the merge, with no one touching anything after the merge click.
- Calendar (#2). An 18:00 booking landed at 23:30, because the server runs in UTC while the calendar test pinned Pune time. The fixer used the library's own time function and removed the test's time-zone pin.
- Share link (#5). The link pointed at
localhost:8080, becausePUBLIC_URLwas set in tests but not in production. Fixed inservice.yamland closed 13 minutes after the failing roll. - Waitlist (#3). The feature flag was on in tests but off in production. The flag is fixed. Its saved check had been written while the page didn't exist, so the re-check honestly said "needs a person"; a
proof-rechecklabel got it a fresh plan, and it held.
It also got things wrong, and we kept the record. Two checks failed for reasons in the check, not the app, and the fixer opened plausible wrong fixes; a person closed both, and one would have raised the cloud bill. That led to:
- setup steps that can't count as failures;
- scoped text checks, and clicking links rather than rebuilding them;
- a plan review before every run;
- a skeptical fixer that reads the recorded verdict;
- a
proof-rechecklabel to challenge any verdict; - a guard that refuses cost changes, and fixes that would close their own issue.
How we built it
- GitLab Duo Agent Platform: a custom flow (flow registry v1, ambient) with three agents, the examiner, clerk and fixer, using GitLab tools:
gitlab_api_get(commits, then merge requests, thencloses_issues);get_issue,update_issue,create_issue_note;create_commit,create_merge_request;run_command, which calls our runner.
- Trigger: Pipeline events, Run when Passed, so a human merge leads to the deploy and the flow with nobody in the middle.
- Agent configuration:
.gitlab/duo/agent-config.ymlruns flow jobs in our runner image. A setup script prefetches the production deployment's merge requests with the job token. - The runner (
proof-run): Node and Playwright.- It validates every plan against the step vocabulary: no custom JavaScript, no leaving the app's origin.
- It photographs each step and outlines pages for the agent.
- It checks calendar files with an RFC 5545 parser, renders the comment strips, and publishes and replays runs.
- The pipeline: eight CI stages in
.gitlab-ci.yml, covering all nine lifecycle stages. These are all real jobs:- verify: three test jobs (four suites: app, runner, board, guard policy);
- secure: SAST and secret detection;
- package: the runner image to the GitLab registry, the app and site images with Cloud Build to Artifact Registry;
- release: release notes linked to the proof;
- deploy: Cloud Run, configured from
service.yaml; - prove: the re-check after a fix;
- monitor: the nightly light table;
- govern: the guard.
- Google Cloud:
- Cloud Run hosts the app and the proof-sheet site;
- Cloud Storage holds the frames;
- Cloud Build and Artifact Registry build and store images;
- Workload Identity Federation with GitLab's OIDC ID token lets CI deploy without a single stored key.
- The proof-sheet site:
- built with Next.js, designed as a photographer's contact sheet:
- real frames on film strips;
- a graphite grease-pencil circle on the frame that proves a criterion held, a red cross on the step that failed;
- a loupe to inspect any frame.
- its API asks GitLab who is calling: a job token must belong to a running job of this project, and a flow token to a member with Developer access.
- CI holds no secrets.
- Hackathon participants have the Developer role, so the project can't mint a bot token. The board therefore holds one delegated token in Secret Manager and acts for CI jobs only when its own stored record backs the action: it re-checks the fix policy on the merge request itself, and closes an issue only if a roll the same job published says it held.
- Tests: runner (19), site and guard policy (17), app (20). All run in CI on every pipeline.
Challenges we ran into
- Developer role only. Participants can't create project access tokens or CI variables. Writes from plain CI (merging a fix, closing after a re-check, reopening a regression) go through the board, which acts only on evidence it stored.
- GitLab's own closing keywords. The fixer's first commit message, "Fix #3: …", closed the issue on merge, before production was re-checked. The fixer now writes "Repair #3 in production", and the guard refuses any fix that would close its own issue.
- Scoped labels are exclusive. A
proof::recheckchallenge was silently dropped when a status label was added, so the challenge label is unscoped (proof-recheck). - Flows can't trigger flows, and bots can't trigger flows. So the loop after a fix runs in plain CI: the guard merges the fix and the re-check job replays the saved plan. This turned into a strength. Replays need no model, so nightly checks are cheap, repeatable and comparable.
- The flow's token can't read deployments. It doesn't have the scope. We prefetch the deployment's merge requests in the agent configuration's setup script with the job token, and fall back to the merge requests on the commit and the issues they close.
- Flow prompts are Jinja templates. Plans therefore use
[[name]]for remembered values;{{name}}would be swallowed. - Cloud Run reserves
/healthz. The version check moved to/health. - Making a model's output safe to run against production. The answer was a small vocabulary validated in code, an origin lock, and verdicts that only the runner's assertions can give.
Accomplishments that we're proud of
- A full hands-off loop: merge, deploy, prove, reopen, fix, guard, merge, re-check, close. The only human action is the first merge.
- Evidence where the team already works. The failing frame is inline in the GitLab issue, and every roll can be opened and verified.
- All nine DevSecOps stages, each one a real job or agent you can click in the pipeline history.
- A tool that says what it can't check. Three outcomes, never two.
What we learned
- The most valuable check is the one CI can't run: production with production's configuration.
- Agents are good at reading intent from prose and writing structured plans. Deterministic code should judge the results.
- Governance is what makes hands-off acceptable: a written policy, a narrow vocabulary, and a tamper-evident record.
What's next
- Canary-first proofs: check a Cloud Run canary revision before shifting traffic, and roll back on failure.
- More checkable criteria: email through a test inbox, API responses, accessibility checks.
- A Catalog-ready template: one flow, one agent config and one CI include, so any GitLab project gets proven deploys in minutes.
- Carbon-aware nightly scheduling: run the light table in the cleanest hour of the grid.
Built With
- artifact-registry
- gitlab-ci
- gitlab-duo-agent-platform
- google-cloud
- google-cloud-build
- google-cloud-run
- next.js
- node.js
- react-js
- typescript
- workload-identity-federation
Log in or sign up for Devpost to join the conversation.