Inspiration
Breaking news does not pause because captions do. Neither does a severe-weather warning, a live game, or the moment in a story everyone talks about next. For Deaf and hard-of-hearing viewers, captions make those moments accessible.
The FCC's television captioning rules define quality through accuracy, synchronicity, completeness, and placement. Caption access is not optional; it is part of delivering the program. FCC closed-captioning guidance
Networks maintain that access across many channels while control rooms operate by exception: health signals tell one operator where to look. But when video, audio, and the overall feed remain live, a caption layer that stops advancing can hide behind a wall of green.
Inside the control room, the broadcast looks healthy. For a viewer relying on captions, it has stopped making sense.
We built Changeover for that gap: the difference between a signal that is still on air and a program that is still accessible.
What it does
Changeover protects caption continuity across multiple live television channels:
- Detects viewer-visible failure. It recognizes when captions stop advancing while video, audio, and the overall feed remain live.
- Establishes whether the evidence is trustworthy. Caption-sync and feed-liveness measurements must be fresh, complete, and matched to the affected channel. Otherwise, Changeover warns the operator and refuses to recommend a switch.
- Diagnoses the failed layer. When the evidence is sufficient, it isolates the caption failure and presents the supporting measurements.
- Verifies recovery. It independently confirms that a healthy backup source is available before proposing a change.
- Resolves contention. If two channels fail while only one backup is available, deterministic policy applies the network’s predeclared service tiers and makes the unresolved consequence explicit.
- Preserves human authority. Changeover can investigate, recommend, or refuse. Only the operator can authorize a feed change.
- Records the decision. After an authorized recovery, Changeover writes the evidence, operator decision, measured outcome, and any unresolved channel to Grafana, then reads the record back to verify the operational trail.
For the operator, this is not another open-ended dashboard investigation. It is a bounded decision with evidence, a clear human gate, and an honest terminal state: RESTORED, REFUSED, or PARTIALLY MITIGATED.
For viewers, access is restored wherever the evidence and available capacity justify it—and unresolved harm never disappears behind a healthy-looking signal.
How Gemini and Grafana power Changeover
Grafana brackets the operational decision: it supplies current evidence before action and retains the verified record afterward. Gemini operates inside that boundary as a bounded diagnostician.
| Point | Sponsor role | Product consequence |
|---|---|---|
| Live channel evidence | Grafana Cloud + official mcp-grafana: Changeover retrieves current, channel-scoped caption-sync and feed-liveness measurements from Prometheus through the official MCP server. |
The decision begins with operational evidence rather than the appearance of the video player. Both signals must be fresh, complete, and associated with the affected channel. |
| Structured layer diagnosis | Gemini 2.5 Flash + Google ADK: After the evidence gate passes, a named ADK Agent and Runner convert the measurements into a structured failed-layer diagnosis. | The operator receives a consistent assessment that downstream code can validate and display. Gemini is never invoked when the available evidence is stale or incomplete. |
| Post-incident audit | Grafana Cloud + official mcp-grafana: After an operator-authorized recovery, Changeover writes a Grafana annotation containing the evidence, authorizer, selected action, measured outcome, terminal state, and any unresolved channel, then reads it back through MCP. |
Engineering can later reconstruct what justified the decision, who authorized it, whether access returned, and what remained unresolved. |
Grafana supplies and retains operational evidence. Gemini diagnoses; ffprobe verifies the backup; deterministic code qualifies evidence and applies capacity policy; the human operator authorizes. When the evidence is not trustworthy, the workflow stops before diagnosis, switching, or write-back.
How we built it
Changeover runs one bounded incident path, with each responsibility separated so a model output can never become a broadcast action by itself:
- Detect. Python identifies captions that stop advancing while video, audio, and the overall feed remain live.
- Collect. The official
mcp-grafanaserver retrieves channel-scoped caption-sync and feed-liveness measurements from Grafana Cloud Prometheus over stdio. Python owns deterministic retries. - Qualify. An evidence gate requires both measurements to be fresh, complete, and associated with the affected channel. Missing, stale, or mismatched evidence ends the workflow as REFUSED before Gemini, backup verification, authorization, or switching can run.
- Diagnose. After the evidence gate passes, Gemini 2.5 Flash runs through a Google ADK Agent and Runner, converting the measurements into a structured failed-layer diagnosis. Gemini cannot switch a feed.
- Verify. ffprobe independently confirms that the proposed backup media is usable.
- Prioritize. A deterministic
ContentionSupervisorpreserves channel identity, enforces available capacity, and applies predeclared service tiers. With two failures and one backup, it proposes the emergency-tier channel and leaves the other visibly degraded. - Authorize. The operator alone approves failover. Changeover waits indefinitely rather than converting a diagnosis or policy result into an autonomous broadcast action.
- Record. After an authorized run, Changeover measures the outcome, writes the evidence and decision to a Grafana annotation through the official MCP server, and reads the annotation back to verify the operational record.
Challenges we ran into
Reproducing the failure reliably. A browser subtitle track was not enough. We authored and timed cues for two films, then built state that can freeze an unfinished line while video continues. Seeking, reloads, and simultaneous channel playback all had to reproduce the same incident.
Treating telemetry as evidence rather than merely data. A caption measurement is not useful if it is stale, incomplete, or associated with the wrong channel. We built an explicit evidence gate that checks both required signals before Gemini or any recovery step can run. Proving refusal required verifying the absence of downstream behavior, not just displaying a warning.
Making the official Grafana MCP integration portable and bounded.
mcp-grafanais an external Go binary communicating over stdio, with different installation paths across macOS, Linux, and CI. We implemented binary discovery and careful process lifecycle handling, then limited the application’s write path to post-authorization annotations with duplicate prevention and MCP read-back verification.Proving the authority boundary. State transitions cross asynchronous UI and backend boundaries. Our tests had to establish that no feed changes before operator approval—and that stale evidence triggers no Gemini diagnosis, backup verification, authorization control, simulated switch, or Grafana write-back.
Accomplishments that we're proud of
Three inspectable operational outcomes. Changeover demonstrates successful recovery when evidence and capacity are sufficient, explicit refusal when the current condition cannot be established, and partial mitigation when one backup cannot restore two affected channels.
A complete Grafana evidence-to-audit loop. The official
mcp-grafanaserver retrieves both required Prometheus signals before the decision. After an authorized recovery, it writes the evidence and outcome to a Grafana annotation and reads the record back for verification.A tested, reproducible build. The project includes unit, sponsor-path, refusal-invariant, and Playwright coverage across the recovery, refusal, and capacity-contention workflows.
A properly attributed media pipeline. We integrated Creative Commons footage from Blender Foundation’s Tears of Steel and Sintel, included the required attribution, and authored timed captions for each scenario.
What we learned
Recovery is an allocation problem, not only a diagnosis problem. Redundant capacity is finite. Restoring one channel can mean leaving another audience without access, so service policy and the resulting tradeoff must remain explicit.
Operational telemetry needs an evidence contract. Retrieving a measurement is not enough. The system must establish that every required signal is fresh, complete, and associated with the correct channel before using it to justify intervention.
Trustworthy automation includes restraint and accountability. Changeover must know when not to recommend action, preserve the operator’s authority when action is justified, and leave engineering with a verified record of what happened afterward.
What's next
The current prototype simulates the final feed switch. Next:
Connect real actuation. Add an SDI/IP router or cloud playout integration behind the existing authorization gate, keeping diagnosis and policy separate from execution.
Pilot with practitioners. Work with broadcast operations and accessibility teams to validate alert thresholds, service-tier rules, evidence presentation, and the approval workflow under realistic incident conditions.
Extend carefully. After the caption path is validated, evaluate sign-language and audio-description services with separate signals, safeguards, and operator validation for each.
Log in or sign up for Devpost to join the conversation.