What inspired me

Anyone who's been on call knows the feeling: the pager goes off at 3 a.m., the dashboard is screaming, and you've got maybe ninety seconds to figure out whether this incident is the one you're allowed to fix — or the one that takes the whole region down if you guess wrong.

I wanted to see what happens when an AI crew holds that pager. Not a chatbot that talks about fixing things — real agents wired into a live production system, watching real metrics, calling a real control plane, and — crucially — being told no when they're not allowed to act.

The question that drove the whole project: How do you give an AI autonomy without giving it dangerous agency? At some point a self-healing system becomes a self-destructive one. I wanted the line to be explicit, enforced, and not decided by the model.

What I built

Director's Cut is a multi-agent broadcast-ops crew. The pipeline is deliberately real:

  • A real ffmpeg encoder streams RTMP through toxiproxy, whose Prometheus metrics are derived from actual ffmpeg -progress output — starved frames vs. expected. Nothing is scripted counters.
  • Grafana is the observability spine, and a self-hosted Grafana MCP server is the agents' only window into metrics and logs. No agent code touches Prometheus or Loki directly.
  • The Director orchestrator delegates through a fixed pipeline: detect → diagnose → policy-check → act-or-escalate → document.
  • The Root-Cause agent pulls live Prometheus range queries + Loki logs + runbook RAG through MCP to produce an evidence-backed root cause with confidence.
  • The Studio Head gate runs a pure-Python, deterministic IAM policy (Cloud IAM-backed) — no LLM in the authorization path, ever.
  • The Remediation agent executes only from a fixed catalog via Cloud Tasks/Pub/Sub, and annotates the Grafana dashboard so the demo shows the agent acting.
  • When policy says Tier 3, it does nothing — and writes a fully pre-filled escalation ticket instead.

The alive demo

I throttled the encoder network to 5 KB/s. Real frames started starving, and the drop rate spiked to 29.6 real drops/sec. The crew caught the alert in Grafana, correlated metrics and logs, checked policy, took the Tier-1 "requeue" action, and the drop rate returned to 0 — then it annotated the dashboard and wrote the post-mortem. Fully hands-off. I broke it; the agents (within guardrails) fixed it.

The "money moment" is governance:

$$ P(\text{act}) = \begin{cases} 1 & \text{Tier 1: reversible, low blast radius} \ 0 & \text{Tier 2: needs human approval} \ 0 & \text{Tier 3: irreversible / global blast radius} \rightarrow \text{escalate} \end{cases} $$

And the actual detail that mattered: a Tier-3 CDN wide-action escalates into a ticket rather than a button push.

How I built it

  • ADK agents (LlmAgent) defined declaratively, deployed to Vertex AI Agent Engine.
  • Live-run measurements came from the running encoder stack, not mocked numbers — real ffmpeg drops, real Prometheus series, real Grafana annotations.
  • The whole thing is validated by a 16-test suite (e2e + policy), and the Agent Engine model call receipts (prompt/thought/candidate tokens) are the proof the LLM actually ran.

What I learned

  • The model is not the guardrail. Authorize with deterministic policy, not vibes. The most important line of code rejects a remediation the LLM never should have been asked for.
  • The demo is the product. Observability tools that show the agent acting on a dashboard beat any slide. Seeing "requeued" appear next to the drop-rate graph is the whole point.
  • Real over simulated. A live ffmpeg encoder failing under real network throttling is worth a thousand fabricated graphs.

Built With

  • adk
  • agent
  • ai
  • bigquery
  • cloud
  • docker
  • engine
  • ffmpeg
  • gemini
  • grafana
  • iam
  • loki
  • mcp
  • prometheus
  • pub/sub
  • rtmp
  • run
  • search
  • tasks
  • toxiproxy
  • vertex
Share this project:

Updates

Submission history