Inspiration

It started with a production incident we never wanted to experience again.

A bug slipped through our deployment process and deleted customer data. What followed was a full-scale war room: engineers, DevOps specialists and database administrators on an emergency call, investigating the failure, containing the damage, recovering data from backups and verifying its integrity.

We already had source control, CI/CD pipelines, monitoring, deployment history and backups. The evidence and the recovery mechanisms existed, but people still had to connect everything under pressure. Most of the room's time went on two questions: what changed, and did the fix work?

The second question is the one that gets skipped. A rollback job turns green, everybody relaxes, and nobody measured whether the service is healthy again.

RollbackIQ is built around that gap. It does not restore deleted data, and it says so. It handles the case that comes first and most often: a release that makes the service slower or makes it fail, and a rollback that must be decided by a person and then proven.

Detect. Decide. Recover.

We built it for GitLab Transcend's Life After Code challenge, as a Path A: Start Fresh project, because software engineering doesn't end when code is merged. Some of its most important work begins after deployment.


What it does

RollbackIQ is a GitLab-native reliability workflow that takes a release from deployment regression to verified recovery. Instead of stopping at an alert or an incident summary, it closes the recovery loop:

  1. Deploy and baseline. A GitLab CI/CD pipeline builds one image, deploys the healthy release and measures a baseline from real HTTP requests.
  2. Detect. When a faulty release is deployed, a detector compares new measurement windows (latency and error rate) with the baseline. It opens an incident only after a sustained breach on two windows. One bad window is not enough.
  3. Propose. The incident carries one rollback proposal with a hash, the target release and the risks.
  4. Decide. A named person approves that exact proposal hash in a manual pipeline job. The approval expires after 15 minutes.
  5. Gate. Before anything changes, a safety gate runs 18 checks: environment allowlist, known-good target, approver identity, approval scope and age, idempotency key. One failed check refuses the rollback. The image digest is then compared with the digest that was verified.
  6. Roll back. The pipeline job deploys the target release.
  7. Verify. A separate job deploys the target again and measures a new window that starts after the rollback. Recovery is declared only from that window. A window that is too short is refused.
  8. Investigate. A custom GitLab Duo flow, started by a mention on the incident issue, reads the evidence and posts supported and contradicted hypotheses, open questions and a recommendation.
  9. Report. An incident console shows the timeline, the approval, the 18 checks and the verification, read from the pipeline's own artifacts.

A successful rollback pipeline does not necessarily mean a successful recovery. Step 7 is the feature that defines RollbackIQ.

RollbackIQ does not treat production changes as harmless AI actions. The agent never changes a deployment. It has no tool that can deploy, approve or roll back.

RollbackIQ is not intended to replace experienced production engineers. It removes repetitive investigation and coordination work so engineers can focus on the decisions that need human judgment.


How we used GitLab

The information needed to investigate a production incident is already spread across the software delivery lifecycle, and GitLab connects it. GitLab is the system of record here for source code, merge requests, pipelines, deployments, approvals and incident evidence.

  • GitLab CI/CD carries the whole lifecycle in one pipeline: an attribution guard, the test suite, the image build, the health-gated deployment, the regression, the approval, the gated rollback and the independent verification. Manual jobs are the human decision points. Evidence travels between jobs as artifacts.
  • GitLab Duo Agent Platform: a custom agent and a custom flow, defined in the repository (duo/), registered in the AI Catalog by a script, enabled in the project, and started by a flow trigger when its service account is mentioned on an issue.
  • GitLab Duo Code Review reviewed the merge requests of this project. Its first review found three real defects, which were fixed; the re-review was clean.
  • GitLab SAST and secret detection run in every pipeline. The first SAST run found a real defect in our own probe, which is fixed; no secret was found.
  • GitLab issues hold the incident record, with links to the pipeline and to the Duo session.
  • GitLab Observability: the workload exports OpenTelemetry traces to the group's collector.
  • Container registry and environments hold the image and the demo-staging target.

DevSecOps stages touched by real workflows: plan (incident issue), code review, build, test, security scanning (SAST and secret detection), compliance check, release, deploy, monitor, incident response and retrospective.


How we built it

Python and FastAPI for the detector, the safety gate, the verifier and the read-only incident API. React, TypeScript and Tailwind for the console. Docker for the demo workload: one image, where a release is a set of deployment inputs. GitLab CI/CD for the lifecycle, GitLab Duo Agent Platform for the investigation, OpenTelemetry for traces.

The central scenario uses two releases. The first establishes a healthy baseline; the second introduces a controlled regression in latency and reliability. The isolated Docker target makes it possible to demonstrate recovery without risking real customer systems.

The same scripts run in the pipeline jobs and on one host, so the lifecycle can be reproduced with one command in about 30 seconds. The scenario is intentionally reproducible so the automation can be tested rather than merely described.


Results we measured

  • In GitLab, pipeline 2932808602, on shared runners, every job green: the faulty release measured at 8.59 and 8.75 times the baseline p95 with 13.3% and 8.3% errors; the gate passed 18 of 18 checks; the verification window had 200 requests over 6.01 s, p95 at 1.014 times the baseline, and no errors.
  • A refusal kept as evidence: on an earlier run the verifier refused to declare recovery, because its window was 1.1 s long and the minimum is 3.0 s. The pipeline failed closed.
  • Repeated runs on one host, real containers: 5 of 5 lifecycle runs recovered and were verified; 5 healthy control runs gave 0 false positives.
  • Six refusals, each by a different guard: no approval, a non-human approver, a wrong environment, a target that was never observed healthy, a wrong image digest, and a replay.
  • 114 automated tests. The test:pytest job runs them in every pipeline; its log in pipeline 2934643614 ends with 114 passed.
  • Security scanners: 0 secrets found. SAST reported 6 findings on its first run; one was a real defect (the probe accepted a URL that is not HTTP) and is fixed, and the others are reviewed and listed in docs/SECURITY.md.

The most useful thing the agent did

The Duo flow's first complete investigation was not flattering. It refused to accept "recovery verified" from the incident issue, because it could not retrieve the verification artifact, and it listed three values in our own incident record that did not match the artifacts. It was right: the three values had been copied wrongly. We corrected the record. It also recommended HOLD and ESCALATE, because nothing in the evidence shows whether the failed requests damaged data. An agent that disagrees with its operators on the evidence is the behaviour we wanted.


Challenges we ran into

The biggest design challenge is the difference between an AI agent that suggests a fix and a system that can safely execute one. Production recovery demands more than convincing natural-language reasoning. A deployment shortly before an incident does not prove it caused the incident, an incorrect rollback can make an incident worse, and a green pipeline is not evidence that users are receiving a healthy service. That is why the agent investigates and reports uncertainty, a person decides, the gate can refuse, and the verifier measures.

The concrete problems we hit:

  • A catalog update that did nothing. Updating a flow in the AI Catalog returned success but kept serving the old definition until the update carried a version bump and a release. The provisioning script now reads the stored definition back and compares it by hash.
  • The headless CLI cannot select a custom flow. We traced this to the CLI source, recorded it, and used the flow trigger, which is the supported path.
  • A recovery window that proved nothing. The first probe finished in 1.1 s. The verifier was right to refuse it; the measurement was wrong. The probe now paces itself.
  • A test that passed locally and failed in CI. Adding the test stage showed that the CI image lacked one dependency. The test was fixed at the cause, not skipped.

Accomplishments that we're proud of

The system refuses more often than it acts, and the refusals are in the repository as evidence. Every number in the video and in this text can be traced to an artifact.

Rather than optimizing for faster alerts or better incident summaries, we designed around one outcome: can we demonstrate that the service has recovered?

  • Safety by design: a named human approves every rollback; the agent has no tool to act.
  • Evidence-based decisions: proposals and investigations are tied to deployment and measurement data.
  • End-to-end accountability: every important action is connected to an auditable GitLab record.
  • Measurable outcomes: recovery is judged on application health, not agent confidence.
  • Reproducibility: a controlled regression scenario that can be run repeatedly.

What we learned

Incident response is not primarily an information problem. It is a coordination and decision-making problem. Teams already have logs, metrics, traces, commit histories, deployment records, test results and security findings. During an incident, someone still has to determine which evidence matters, what action is safe, who must approve it, and whether the system has recovered.

Trustworthy autonomy requires clear boundaries. An agent can investigate independently, but consequential production actions must be constrained by evidence, permissions and risk. The strongest AI systems will not be the ones that execute the most actions. They will be the ones that know when to act, when to ask for approval, and how to verify the result.

The definition of done for incident automation is restored service health, not a completed automation workflow.


Limits, stated plainly

  • The fault is injected deployment configuration on one image. It is not a code regression in a different binary.
  • A code rollback does not restore deleted data. Data recovery is a separate, human-controlled operation that this project does not perform.
  • One environment (demo-staging), one fault setting, small request volumes. Five runs are too few to state a rate.
  • In the recorded pipeline, the approval job was played through the API by the token owner. The gate enforces a real, non-agent identity; the recording does not show a person clicking.
  • The Duo flow investigates and recommends. It does not run the lifecycle.
  • GitLab MCP is not used. DAST, dependency scanning and container scanning are not enabled. No Google Cloud deployment is claimed.

What's next for RollbackIQ - Detect. Decide. Recover.

Next steps: dependency and container scanning, an approval through GitLab protected environments, a fault near the detection thresholds to measure false negatives, and a second service to test the design beyond one workload.

Beyond the hackathon:

  • More recovery strategies: feature-flag changes, canary traffic adjustments and configuration restoration, each behind the same approval and gate.
  • Database-aware recovery: checks for schema compatibility, migration reversibility and data integrity. Our inspiration involved deleted customer data, so this matters. Future versions could coordinate validated backup recovery procedures while keeping destructive or irreversible operations under strict human control.
  • Predictive release risk: use historical deployment outcomes and incident patterns to identify high-risk releases before they reach production.
  • Organizational incident memory: learn from previous incidents, approved decisions and verified outcomes to improve future investigations.
  • Sustainability-aware operations: investigate whether earlier detection and fewer unnecessary pipeline reruns reduce wasted compute. Any environmental claim will be based on measured data and a documented method.

We want RollbackIQ to become a safety layer between software deployment and production reliability: a system that can observe what changed, understand what went wrong, coordinate a safe response, and verify that customers are no longer affected.

Because the best production incident is the one that never becomes a war room.

RollbackIQ — Detect. Decide. Recover.


Built With

  • agentic-ai
  • ai-agents
  • gitlab
  • gitlab-api
  • gitlab-ci/cd
  • gitlab-duo
  • gitlab-duo-agent-platform
  • gitlab-duo-cli
  • gitlab-mcp
  • incident-response
  • multi-agent-systems
  • rest-api
  • tailwind-css
Share this project:

Updates

Submission history