Inspiration
All tests passed. Five videos were still wrong.
A media pipeline can install an update, pass its software tests, and produce valid MP4 files while quietly moving captions outside the delivery safe area. The failure only becomes obvious when someone looks at what the client will receive.
I built Greenlight around that gap: how can a pipeline engineer approve an update with evidence about the actual output?
The demo follows Avery, a fictional media platform lead with a campaign waiting to render. Her question is practical: should she merge this compositor update before the campaign runs?
What it does
Greenlight tests a pipeline change against a small, reproducible media canary, brings the evidence into a Grafana investigation, and produces a release Decision Card.
- Render both versions. Eight original synthetic clips become sixteen baseline and candidate outputs.
- Check the pixels. Greenlight measures caption bounds in the raster and separately checks three decoded frames from each final MP4, alongside file dimensions and duration.
- Investigate through Grafana. A Gemini agent executes a bounded investigation through Grafana MCP, bringing together the alert, Prometheus metrics, Loki logs, and Tempo traces.
- Apply the policy. The deterministic evaluator returns
HOLDwhen the candidate violates the delivery contract. The agent explains and annotates the evidence; it cannot change the thresholds or override the verdict. - Repair and replay. The corrected compositor runs against the same clips and policy. With the caption failures resolved, Greenlight returns
PROMOTEfor human review.
The Decision Card puts the visual comparison, failing clip, reasons, and next action together. Avery can keep the current version, correct the candidate, and repeat the check.
How we built it
The core is TypeScript and Node.js, with FFmpeg generating original synthetic media and FFprobe checking the encoded files. The web interface is a lightweight Decision Card, hosted on Google Cloud Run.
Google ADK runs Gemini on Vertex AI. Its Grafana MCP integration retrieves investigation evidence and writes an annotation linking the affected clip, trace, and decision receipt. OpenTelemetry carries the experiment's metrics, logs, and traces into Grafana Cloud.
The receipt binds the decision to its inputs, policy, and evidence. Google Cloud KMS signs the recorded live receipt, which can be verified offline with an exported public key.
The hosted app displays a prepared result. The demo includes the real September 3 Grafana investigation and a separate corrected local run. Synthetic fixtures are labelled throughout. The current agent follows a guided investigation plan; independent diagnosis is a next step.
Challenges we ran into
Checking the deliverable accurately. Video compression introduces pixel differences of its own. Checking decoded MP4 samples required a matched encoded control and compression-aware filtering, while preserving the exact raster measurement.
Giving the agent useful tools with clear limits. Tool names, arguments, and scope must be authorized before execution. Detecting an unauthorized operation afterward is too late.
Keeping evidence trustworthy. A successful tool call does not automatically prove a measurement. Greenlight separates evidence retrieval from operational actions such as annotations, checks returned log values against local measurements, and reports incomplete coverage.
Making repair meaningful. The same pipeline had to recognize a corrected candidate, with unchanged thresholds, rather than only reproduce the known failure.
Accomplishments that we're proud of
- Reproduced five failing candidate clips, including a measured 61-pixel caption overflow, while every baseline passed.
- Completed a real eight-call Grafana MCP investigation with Gemini and Google ADK.
- Produced a signed live decision receipt with offline verification.
- Demonstrated
HOLDfollowed by a corrected localPROMOTE, using the same eight clips and policy. - Built a reproducible local path with tests covering policy authority, media failures, and tool authorization.
What we learned
A valid file is not necessarily a valid delivery. Checking output pixels catches failures that container checks and software tests can miss.
We also learned that an agent does not need authority over acceptance criteria to be useful. It can help assemble and explain operational evidence while a deterministic policy remains responsible for the release verdict.
What's next for Greenlight
The next step is a bounded investigation in which the agent identifies the affected clip without receiving the known failure in advance.
From there, I want to connect a real post-production queue, expand beyond static caption fixtures to changing captions and additional delivery checks, and measure investigation time with pipeline engineers. The aim is to make update reviews faster while keeping every release decision reproducible.
Built With
- cloudkms
- cloudrun
- ffmpeg
- gcp
- gemini
- googleadk
- grafana
- loki
- mcp
- opentelemetry
- prometheus
- tempo
- typescript

Log in or sign up for Devpost to join the conversation.