Inspiration
When a live sports stream or a premiere starts buffering, the repair is usually fast — fail over a CDN, roll back an encoder, drain a bad region. What's slow is figuring out what is wrong. Industry benchmarks put the average media incident at ~30 minutes to detect plus ~40 minutes to resolve, and most of that hour-plus is a human doing swivel-chair correlation across metrics, logs, and traces — by CDN, by region, by device — under pressure, while media & entertainment downtime runs up to $2M per hour. We wanted to automate the part that actually hurts: the diagnosis, not the fix.
What it does
Signal watches a streaming platform through Grafana and, when quality-of-experience degrades, runs a fixed, deterministic 7-step incident-response playbook:
Detect — polls QoE SLOs via PromQL (rebuffer ratio, p95 start time, error rate). Triage — classifies severity and blast radius by CDN/region. Correlate — pulls Loki logs for the failing segment. Change-check — looks for recent deploy annotations. Diagnose — Gemini ranks root-cause hypotheses, each citing the specific signals. Recommend — links the scoped Grafana dashboard and proposes a mitigation. Report — produces an engineer-grade RCA and a one-paragraph exec summary.
In our live demo it detects a real Akamai / us-east incident — rebuffer ratio, start time, and error rate all breaching SLO simultaneously — isolates it by comparing against the healthy CDNs and regions, and recommends failing over Akamai in us-east. All of it runs from a public URL a judge can click.
How we built it
Gemini on Vertex AI (Gemini Enterprise Agent Platform) does the reasoning, orchestrated with the Google Agent Development Kit (ADK) as a deterministic SequentialAgent — the model's judgment is deliberately confined to the diagnose and report steps so the flow is reproducible. The Grafana MCP integration is the heart of it: we register the open-source mcp-grafana server (65 tools — PromQL, LogQL, dashboards, IRM) as an ADK MCPToolset over stdio, authenticated with a Grafana service-account token, and call it at runtime. Everything is hosted on Google Cloud: the agent runs as a FastAPI app on Cloud Run under a least-privilege service account, with the Grafana token in Secret Manager. A second always-on Cloud Run service generates synthetic streaming metrics and remote-writes them to Grafana Cloud via Grafana Alloy, keeping a live incident present so the hosted demo works with no local machine involved.
Challenges we ran into
The hosted mcp.grafana.com endpoint uses an interactive OAuth 2.1 browser flow — unusable for a headless agent — so we pivoted to the self-hosted mcp-grafana server with token auth, which is also more reproducible for judges. A subtle one: naming our Python package signal shadowed the standard-library signal module, which anyio (deep under the ADK) imports — every run crashed until we renamed it. google-adk gates its MCP tools behind the mcp package, and mcp 2.0 broke the ADK 2.6 API — pinning mcp<2 fixed the tool discovery. On Vertex, a stray GOOGLE_APPLICATION_CREDENTIALS in the shell profile silently forced a service account from an unrelated project, causing 403s that no amount of re-login fixed until we traced it via the token's identity. The Gemini API free tier's 5-requests-per-minute limit couldn't sustain a 7-agent pipeline, which is what pushed us to Vertex AI — the right home for it anyway. Accomplishments we're proud of A genuinely deterministic multi-step agent that produces a sourced, correctly-scoped root-cause report against real telemetry — not a canned demo. The whole thing is self-hosted and always-on in the cloud: a judge can open the URL days later, with our machines off, and still watch it detect a live incident. The agent's honest reasoning — in one run it tried a Loki correlation, self-corrected its LogQL, and reported that no logs matched — reads like a real SRE, not a hallucination.
What we learned
Confining the model to specific steps of a deterministic workflow — rather than letting it free-run — is what makes an agent trustworthy enough to put in front of an on-call engineer. And the highest-leverage automation in incident response isn't the remediation; it's collapsing the time-to-diagnosis.
What's next for Signal
Wire the recommend step to actually execute mitigations through Grafana IRM and CDN APIs (with human approval); add Grafana AI Observability to trace the agent itself; push synthetic logs to Loki so the correlate step surfaces real evidence; and learn per-service SLO thresholds instead of static ones.
Built With
- cloud-run
- gemini
- google-adk
- google-cloud
- grafana
- grafana-alloy
- grafana-cloud
- loki
- model-context-protocol
- prometheus
- python
- secret-manager
- vertex-ai

Log in or sign up for Devpost to join the conversation.