Inspiration

Most pipelines run every job on every change. Fix one line of docs and you still pay for tests, scans and an image build. We wanted a lane that runs what the change needs, lets an agent check that decision instead of trusting it, and says plainly what each run cost.

What it does

GreenLane plans the pipeline from the change.

  1. plan reads the files a merge request touched and writes a child pipeline with only the jobs those files need. A one-line docs change ran 2 of 11 jobs; a change to the service ran 7 of 11; a change to the lane's own rules runs everything, and the plan says why.
  2. The child pipeline tests, scans, builds the image, writes a bill of materials and scans the image.
  3. Its last job writes a carbon note: which jobs ran and why, which were skipped, and the estimated grams of CO2e against a run-everything baseline, with the lane's own two jobs counted against it.
  4. Lane Steward, a custom GitLab Duo flow, is the second pair of eyes. The planner only looks at file paths; the agent reads the diff. Started by a trigger when a merge request pipeline passes (or by a mention), it checks the skipped jobs against the changed lines, explains a failed job from its log, and posts one note with the verdict and the carbon lines quoted exactly.
  5. Lane Mechanic, a second Duo flow, repairs. Started by a trigger when a merge request pipeline fails (or by a mention), it reads the failed job's log and the code. If the log shows the cause directly and the fix is a few lines in a file the merge request changed, it makes one commit on that branch and says what was wrong; otherwise it changes nothing and says what it found. It may not edit a test, the lane's configuration or the flows, and it does not merge.
  6. On the default branch the lane deploys to Cloud Run with no stored key: the job's ID token is exchanged for a 15-minute Google token. It then checks that the new version answers and returns to the previous revision if it does not.
  7. Every two hours a scheduled pipeline runs one job, monitor: real requests to the deployed service (health, a conversion, the page), their latency, and the version it reports against the default branch's head, kept as an artifact. A service that is down or answers wrongly fails the pipeline, so the pipeline history is also the uptime record.

How we built it

  • GitLab CI/CD: a parent pipeline, a dynamic child pipeline, a job catalogue of hidden templates, merge request pipelines, the project's container registry, a CycloneDX report, environments.
  • GitLab Duo Agent Platform: two custom flows defined in the repository (flows/lane-steward.yml, flows/lane-mechanic.yml), each with its own tool list and its own service account. Flow triggers start them from pipeline events on merge requests (failed for the Mechanic, passed for the Steward), filtered to merge request pipelines so child pipelines and the agents' own sessions never set them off; a mention works too. The Mechanic also has create_commit.
  • Google Cloud: Cloud Run, Artifact Registry and Workload Identity Federation with GitLab's OIDC token. No service account key exists anywhere. One instance at most, none when idle.
  • The planner, the pipeline writer and the carbon accounting are Python standard library only, so the planning job installs nothing.
  • Carbon: the Software Carbon Intensity shape (energy times grid intensity, per pipeline), with Cloud Carbon Footprint's coefficients for Google Cloud and Google's 2025 grid intensity per region.

The agents at work (both on record in the project)

  • Merge request 4: the Steward confirmed the plan for a service change and explained why none of the four skipped jobs applied.
  • Merge request 5: we planted a swapped pair of digits in a new unit (3875.41 for 3785.41 ml). The pipeline failed in test-api. Mentioned, the Mechanic committed the one-line fix 56 seconds later, quoted the assertion from the log, explained from the description and the arithmetic why the table was wrong and not the test, and asked a person to check the new pipeline, which passed. The Steward then checked that run's plan. This is one easy case, not a benchmark.
  • Merge request 8: nobody was mentioned. The pipeline failed on an undefined name; the trigger started the Mechanic, which committed the one-line fix two minutes after the failure and said why. When a person's next push passed, the trigger started the Steward, which confirmed the plan and quoted the carbon lines.
  • Merge request 9: the case the Steward exists for. The change adds gal to the page's unit list and nothing else; the plan is right by file paths (2 of 11 jobs) and the pipeline is green, but the API has no gallon. Started by the trigger, the Steward read the API, said the change is "incomplete in a way the path-based planner could not see", named the 400 the page would get, the jobs it would have run and why they would still have passed, and added that the carbon saving "does not mean the change is complete".
  • Merge request 7: a case made to see it refuse. The change rounds results to two decimals, as its description asks, and two untouched tests then fail. Mentioned with the same words, the Mechanic committed nothing: "the code and the tests disagree, and both are defensible"; it would not loosen the tests and said a person should decide the precision.

What the numbers are, and are not

  • A full run of all ten jobs measured 172 seconds of runner time, about a quarter of a gram of CO2e by this estimate.
  • A one-line docs merge request ran 2 of 11 jobs: about 75 % less runner time and carbon, the lane's own jobs included.
  • A change to the service ran 7 of 11: about 23 % less.
  • A full run through the lane costs about 12 % more than a plain pipeline, because planning and accounting are extra jobs. The carbon note says so.
  • Added up over every run so far (the carbon ledger, rebuilt from the public job logs): 39 runs with a measured baseline, 8.17 g CO2e through the lane against 11.14 g for running everything, 2.96 g (27 %) and 3,029 runner-seconds not spent, the lane's own 1.38 g of planning and accounting counted against it. 21 runs skipped something; 18 were full runs that cost more.
  • These are estimates from public averages, not measurements. They leave out embodied emissions and the model inference of the agents. Per pipeline the grams are tiny; the point is the share of runner time that was never spent.

Challenges we ran into

  • A child pipeline under a merge request pipeline came out empty until it carried its own workflow rule.
  • Our first guessed job durations were several times too high and overstated the saving; they were replaced by measured ones.
  • The first run of the agent used GitLab's default template instead of our flow: the definition had not been saved. Flows are now written through the platform's API and read back.
  • The flow's token cannot download job artifacts, so the carbon note is printed in the job log, where the agent can read it.
  • The first deployment never became ready: the service crashed at import inside the container, and nothing in the pipeline had ever started the container. The second served but failed our own check, which asked /healthz, a path Cloud Run answers itself. The third passed. All three pipelines are in the history.

Accomplishments that we're proud of

Every claim above points at a pipeline, a merge request or an agent's note that anyone can open, failures included. The lane also reports the case where it costs more than it saves.

What we learned

A planner that only reads file paths needs a reader of the diff beside it. An agent that may commit needs a short list of things it may never touch more than it needs a long prompt. And a carbon figure is only worth showing with its baseline, its method and its limits next to it.

What's next

Harder cases for the Mechanic and a record of when it declines; the Steward re-planning with a job it would have run; triggers on merge request creation and pipeline failure; GitLab's own security analyzers, so findings reach the vulnerability report.

Built With

  • bandit
  • docker
  • gitlab
  • gitlab-ci
  • gitlab-duo
  • gitleaks
  • google-cloud-run
  • python
  • trivy
  • workload-identity-federation
Share this project:

Updates

Submission history