Inspiration

Writing the code is the easy part now. The hard part is the rest: review, security, tests, deployment, monitoring. That is the work we are handing to agents, and agents are doing it fast. But when an agent reviews a merge request at 2 AM, runs a security scan, and ships a deployment, there is a question nobody in the room can answer: what actually happened? Did the security scan really run, or did the agent say it did and move on? What were the actual test results? Who approved that deployment? Today the answer is: trust me. For a hobby project, trust me is fine. For an enterprise pipeline, trust me is a non-starter. That gap is what we built for. Every agent action in the pipeline should leave proof: captured at execution time, impossible to quietly rewrite, checkable by anyone with a link. Not a memo written after the fact. Evidence.

What it does

Receipted Pipeline is an AI agent that runs five real DevSecOps stages against a real GitLab project, end to end, with no human in the loop: (1) Code review: the agent reviews the merge request and records its findings. (2) Security scan: the agent runs GitLab's SAST and secret detection jobs and captures the results. (3) Testing: the agent executes the test suite and captures pass/fail with details. (4) Deployment: the agent deploys the app to GitLab Pages and captures the deployment record. (5) Monitoring: the agent checks the live health endpoints and captures the status. Each stage produces a verifiable execution receipt (built on AER-1, an IETF Internet-Draft for agent execution receipts). Each receipt is bound to that stage's real GitLab-native evidence: merge request IDs, pipeline and job IDs, commit SHAs, and artifact hashes. When the run finishes, you get a dashboard of receipts, each a public URL anyone can check independently. Honesty note: a verifiable receipt proves the saved execution result was not changed. It does not prove the AI was right. Proof of the event, not proof of correctness.

How we built it

The demo target is a small Flask app with a real GitLab CI pipeline defining five stages: review, security, test, deploy, monitor. Security scanning uses GitLab's own SAST and secret detection templates, which emit real downloadable JSON reports per job. Above the pipeline sits a Python agent orchestrator that walks the stages in order: it reviews the merge request, triggers and reads back the security scan jobs, runs the test suite, pushes the deployment to GitLab Pages, and probes the health endpoints. At each step it pulls the real GitLab objects (MR ID, pipeline and job IDs, artifact hashes) and feeds them into a receipt issued through the Zambo MCP server, formatted to the AER-1 IETF Internet-Draft. The orchestrator is hands-off by design: one command starts the run, and the agent does the rest. The five receipts are served in a single dashboard view, each with its verification URL.

Challenges we ran into

Getting a receipt to mean something real was the central challenge. A receipt of a fictional event is theater. So the binding work, the plumbing that ties each receipt to actual GitLab objects, took the most effort: resolving the right pipeline and job IDs at run time, hashing the real artifact payloads, and making sure the receipt captures evidence the judge can check independently rather than claims the agent makes about itself. The second challenge was orchestration reliability: an agent that stops halfway through a five-stage run produces four receipts and a broken story, so the orchestrator had to be built for hands-off completion. The third was honesty in the copy itself: the receipt proves the record was not rewritten, not that the agent was right, and writing that distinction clearly was a genuine part of the build.

Accomplishments that we're proud of

We covered five full post-code DevSecOps stages in a single hands-off agent run. Each stage's receipt is bound to real, independently checkable GitLab evidence, not agent self-reports. The standard underneath is open and implementation-agnostic: AER-1 receipts work across agent frameworks and CI systems, with a conformance kit anyone can run. The demo answers the judge-facing criterion directly: the video shows the automation running end to end, with the receipts open and verifying live.

What's next

Three things. First, production hardening: extending the stage set beyond the five (dependency scanning with a real lockfile, DAST against the staging environment, rollback-on-failure receipts), and making the orchestrator multi-project aware. Second, integration depth: emitting receipts from inside GitLab CI jobs themselves, so the receipt is part of the pipeline rather than an observer of it. Third, the standards push: AER-1 is a living IETF draft, and every independent implementation that ships makes verifiable execution receipts a default expectation rather than a novelty.

Built With

  • aer-1
  • ai-agents
  • devsecops
  • flask
  • gitlab
  • gitlab-ci
  • hash-chains
  • mcp
  • python
  • sast
  • verifiable-execution-receipts
Share this project:

Updates

Submission history