-
-
CarbonTwin release steward: the lifecycle and the results on two database releases.
-
The live dashboard on Firebase Hosting after release v2026-08.
-
A batch the agent opened and the release steward merged after its pipeline passed (MR !15).
-
The pipeline on main after that merge: tests, security scans, build, deploy, live check, package, release and SCI.
-
The release issue where the agent reported each batch of v2026-08 (MR !13 shows as closed; its batch is on main).
-
Software carbon intensity per pipeline, and the model calls that reusing verdicts avoided.
Inspiration
CarbonTwin (MIT, built for the IEEE ClimateChain hackathon) screens Berkeley's Voluntary Registry Offsets Database for projects listed in two carbon registries that issued credits for the same vintage. Its published run on release v2026-06 found 18 such pairs, covering 2,486,944 tCO₂e on the smaller side.
An audit like this goes stale with the next release. Berkeley republishes the database about every two months, and a release can rename IDs, add descriptions and change the status of listings that were already judged. One-off analyses usually stop there. We wanted the audit to keep itself current without a person in the loop, and without paying for a full re-run each time.
What it does
A daily scheduled pipeline watches the database page, and a second schedule, every 15 minutes, creates a pipeline only while pairs are waiting for a verdict. When a pipeline passes, a pipeline-event trigger starts a custom GitLab Duo Agent Platform flow, the release steward, which takes one step per session:
- Rescreen and carry. Scripts rescreen only the countries whose records changed (identical to a full run) and move cached verdicts to renamed IDs. A verdict stays valid while both listings read the same to the adjudicator: name, developer, location, type, methodology, notes and description. Issuance numbers and registry status do not count, because code computes the overlap and status changes do not change the asset.
- Adjudicate one batch. The agent judges up to 30 pairs whose verdicts are missing or stale, largest overlapping volume first, and writes one JSON verdict per pair with the shared facts it relied on.
- Check the agent. Code validates every line against the schema and looks for each cited fact in both listings. Unsupported facts are dropped; a pair is confirmed only with two grounded facts. The report is rebuilt and its Merkle root anchored.
- Merge only what the pipeline verified. The agent opens a merge request. Its pipeline recomputes the merge on a copy and fails on any difference, and runs the tests, SAST, secret detection and dependency scanning. The next scheduled session reads that pipeline: if it passed, the steward fast-forwards
mainto exactly the tested commit; ifmainmoved in the meantime, it rebases and lets the pipeline run again. - Ship. After the merge, the dashboard deploys to Firebase Hosting on Google Cloud (and to GitLab Pages), a live check reads the public site back and confirms the new Merkle root, the verified data is published to the package registry, and a release links to it.
- Report. The agent notes every batch on a "Database release" issue and closes it when nothing is waiting. CI jobs are measured with CodeCarbon, flow steps record their run time, and each pipeline reports its software carbon intensity.
Results on two real database releases
The repository starts from release v2026-04 so that the steward replays the real v2026-04 to v2026-06 transition. Between those releases, 920 project IDs appeared and 795 disappeared; 784 of them were renames (ACR102 became ACR0102), and 721 of the renamed ACR listings gained a description they did not have before.
- Adjudicate every covered pair again: 450 model calls
- Reuse a verdict only when both full records are unchanged: 330 model calls
- Reuse a verdict while both identity texts are unchanged (CarbonTwin): 237 model calls
On GitLab the steward worked through this release in eight batches on 6 October 2026, between 15:36 and 20:04 UTC. Each batch took two scheduled sessions, one to adjudicate and one to merge, and its merge request stayed open for 6 to 26 minutes. The agent opened seven of the eight merge requests and the steward merged all eight; we opened the first by hand while fixing how the flow creates merge requests. From the second merge on, every merge deployed the dashboard and published a package and a release. See the pipelines, a merge request, the release issue, the package registry and a release.
The agent judged 237 pairs and agreed with the published verdict on 192 (81%). The release sent 23 published findings back to it, because a listing had been renamed or redescribed; it confirmed 22 again and marked one as unclear (same name and developer, different project types). It also confirmed three pairs that the published run's second pass had not, one of them with a 2020 vintage issued in both registries (718 tCO₂e). Of the 45 disagreements, 37 were pairs the published run had judged the same asset but could not ground in both listings, so it had already left them for review. After the replay the report lists 63 same-asset pairs (published: 61), 19 of them with a shared vintage (18), covering 2,487,662 tCO₂e (2,486,944).
Then a real release arrived. Berkeley published v2026-08 late on 6 October (UTC). The steward picked it up at 00:08 UTC the next day, in a session that an unrelated README push happened to start (the daily watch would have found it at 06:00). It kept 318 of the 450 covered verdicts and had the agent judge the other 132, then 13 flagged pairs outside the coverage whose listings had changed (see Challenges): 145 pairs in six batches. The first batch waited for the daily watch to merge it, because the 15-minute clock starts only after a merge has recorded on main that pairs are waiting. Two pairs gained a vintage issued in both registries, among them ACR0931 and CAR2112 for 2025 (147,508 tCO₂e). The report now lists 63 same-asset pairs, 21 of them with a shared vintage, covering 2,762,944 tCO₂e.
A local check, ci/replay_check.sh, replays all eight batches with the published verdicts standing in for the agent. Every batch passes the merge request checks, and the end state contains all 61 findings of the published report with the same facts.
Sustainability
Most of the footprint is model work, so the design avoids it: 237 model calls instead of 450 for the v2026-06 transition (47% fewer), and 145 for v2026-08, where 318 covered verdicts were reused. Scheduled pipelines run only the watch job, because nothing in the repository changed; the 15-minute clock creates no pipeline at all once nothing is waiting, so no agent session starts; and the flow's trigger ignores merge request pipelines.
Each pipeline reports SCI = ((E × I) + M) / R with R = one pipeline run. E comes from CodeCarbon for the CI jobs (flow steps run in GitLab's agent sandbox, where CodeCarbon cannot read process data, so only their run time counts, in the embodied share); model energy per pair is 0.95 Wh (Jegham et al. 2025, short query on a Claude Sonnet model), reported with a 0.24 to 2.99 Wh range; grid intensity 384 gCO₂/kWh (Ember, US 2024); embodied emissions 0.00033 kg per runner hour (Boavizta, 2-vCPU instance).
A merge request pipeline for 30 pairs comes to about 11 g CO₂e (2.8 to 34 g with the low and high values), almost all of it model work; a scheduled or main pipeline stays under 0.01 g. The v2026-06 replay came to about 97 g CO₂e over the 45 pipelines that reported SCI (25 to 307 g), and reusing 213 verdicts avoided about 78 g (20 to 245 g). One overlapping session judged a batch a second time before we added the claim described below; its 30 calls are not in these totals. Release v2026-08 came to about 53 g CO₂e over 25 pipelines (13 to 167 g), and reusing 318 verdicts avoided about 116 g (29 to 365 g).
How we built it
Python (pandas, scikit-learn) for screening and validation, Node.js for the Merkle anchoring and the dashboard, Solidity for the registry contract, GitLab CI/CD, the package registry and releases, and a custom flow on the GitLab Duo Agent Platform running in the default sandboxed image with network access limited to the database host. The dashboard is served publicly from Firebase Hosting on Google Cloud. The pipeline deploys it without a stored key, exchanging GitLab's ID token through Workload Identity Federation, and the Google Cloud project has no billing account, so it stays on the free Spark plan. GitLab Pages in the hackathon group is visible to project members only.
Challenges
- The v2026-06 release renamed 784 IDs and filled in descriptions for the same listings. Keying the cache on full records sent 93 more pairs to the model; keying it on what the adjudicator reads fixed that.
- Pipeline-event triggers fire for every pipeline, including the steward's own merge requests, but GitLab does not start a flow from a pipeline that the flow's own service account created. A second trigger meant to merge the steward's merge requests therefore never fired. The steward now takes one step per session on a schedule: merge the open batch if its pipeline passed, otherwise start the next one.
- Sessions overlap when a scheduled run fires while another session is still judging. A failed branch listing once read as "no batch open" and the same batch was opened twice. The prepare script now stops when it cannot list branches, the batch branch is named after a hash of its pairs, and a session takes a claim branch atomically before any work, so an overlapping session skips instead of judging the same pairs again.
- After the first merges, GitLab refused to create the scheduled pipeline on a batch commit: a release job chosen by the commit message needed the live check, which scheduled pipelines leave out. No pipeline meant no session, and the steward stalled without an error anywhere it was looking. Those jobs now run only in push pipelines, and a test evaluates every job's rules and needs for each pipeline source.
- Release v2026-08 added a 2025 vintage to 34 more pairs, which pushed lower-ranked pairs out of the 450-pair coverage while it changed their listings. Thirteen of them had been judged the same asset, and their findings dropped out of the report without any error. A pair outside the coverage now goes back to the agent when its verdict had flagged it and its listings change; the agent confirmed all thirteen again.
- One merge request showed as closed although its batch was on
main: the merge script deleted the branch right after pushing, before GitLab had registered the merge. The next session now removes merged branches. - The flow's
create_file_with_contentstool refuses to overwrite a file. A leftover answer file made the agent fall back toedit_fileand a shell heredoc and write its 30 verdicts three times; removing the file before each batch brought a session down to eight tool calls. - The two releases differ in column count, and the loader read vintage columns by position. It now finds them by their labels and stops on any layout change; the agent escalates that to an issue for a person.
What was added during the submission period
This is a Path B entry. The original project (MIT, https://github.com/HyunsikParker/carbontwin) is the screening pipeline, the model adjudication, the contract and the dashboard. Everything about the release lifecycle is new, built from October 5, 2026: incremental rescreening, identity-keyed verdict cache and rename carry, batch queue, verdict validation and merge, the Duo flow and its configuration, the CI/CD pipeline, release issue reporting, packaging, the public deployment and its live check, carbon measurement and the replay tooling.
What's next
Keep running on every new Berkeley release; merge a release's first batch sooner (today it waits for the daily watch, because the 15-minute clock starts only after a merge has recorded that pairs are waiting); and add a second adjudicator for pairs where the agent and the published run disagree.
Built With
- codecarbon
- firebase
- gitlab
- gitlab-ci
- gitlab-duo
- gitlab-pages
- google-cloud
- node.js
- pandas
- python
- scikit-learn
- solidity
- vite
Log in or sign up for Devpost to join the conversation.