Inspiration

What it does

How we built it

Challenges we ran into

Accomplishments that we're proud of

What we learned

What's next for Broadcast QC Incident Commander

Inspiration

Automated QC tools already measure everything. Baton, Vidchecker and Venera all produce a report saying the file is out of spec. None of them tells you why, and none tells you who to talk to. The expensive part of a rejected delivery was never detection - it was attribution, inside an SLA window.

The failure worth solving is not the obvious one. Here the normaliser does its job correctly and the defect appears one stage later:

Stage Preset Integrated loudness Verdict
ingest ingest_passthrough_v1 -23.1 LUFS PASS - source arrived in spec
normalize norm_ebu_v3 -23.1 LUFS PASS - normalisation was correct
package pkg_h264_v7 -17.5 LUFS BLOCKED

pkg_h264_v7 applies pan=stereo|c0=c0+c1|c1=c0+c1 - an unconditional downmix that sums an already-stereo bed into both output channels. You cannot guess that from the failing number: the source was fine and the normaliser was fine. The investigation has to walk the stages.

What it does

  1. Runs a three-stage delivery pipeline on real media with real ffmpeg
  2. A deterministic gate - zero AI - measures and blocks against a YAML profile
  3. An agent investigates: it chooses its own read-only Grafana tools through the official Grafana MCP server, follows what the data shows, and forms a hypothesis
  4. It tests that hypothesis by re-running the failing stage with the preset that normally runs, and measuring both. A preset that ran is not a preset that caused
  5. It proposes a repair from a typed allowlist; a human approves
  6. The repair is re-validated by the same gate, then written back to Grafana as an annotation and an IRM incident

Attribution targets a preset version, not a worker. Real facilities do not hunt for a broken machine; they hunt for which transcode preset changed, and when. Provenance rides the trace span and comes back as a cited claim: pkg_h264_v7 v7 was changed by d.okonkwo under CHG-4471, approved by j.reyes. That ticket was properly approved and still shipped a defect - which is precisely why attribution has to be measured from telemetry rather than inferred from process.

It also declines to answer. Select the Netflix profile and the same asset returns UNMEASURABLE, not a verdict: Netflix specifies dialogue-gated BS.1770-1, this probe produces programme-gated BS.1770-3/4. Different selection and different gating, so the two would disagree over identical audio. A confident number that is wrong by an unknown amount is worse than no number at all.

How I built it

The model interprets and proposes. Deterministic code gathers, adjudicates and executes. Three boundaries enforce that:

The controller binds provenance; the model only interprets. EvidenceLedger.observe() is called by the controller after a query runs, binding the phase, the query text, its hash, and a reference to the raw response. The model's entire surface is one tool: record_evidence(finding, supports). It cannot name the phase, the query, or the step id. Without this, a model can emit a perfectly schema-valid evidence record citing a query it never ran, and the integrity story is theatre.

The agent has no execution authority. It names an action_id from a per-profile allowlist and supplies typed parameters; it never emits a command string. Every request is re-validated before running, and repairs always write a new artefact rather than overwriting the input.

The gate is not the model. pipeline/profiles/ebu_r128.yaml is the authority on pass/fail. The same code that blocked the asset clears the repaired one - if the agent were wrong, re-validation would say so.

Signal routing is deliberate. QC results are test records, not operational time-series - pushing per-measurement values into Prometheus keyed by asset_id is a cardinality anti-pattern. Measurements go to Loki; the asset's journey goes to Tempo. Every log line carries trace_id and span_id, without which the traces and logs are two disconnected piles and there is no investigation to run.

Accomplishments I'm proud of

Four adversarial refusals, run through the same validator as the real path, so the refusal is reproducible rather than dependent on the model misbehaving:

Candidate Refused because
fabricated_citation cites unknown step step-99
laundered_citation one real citation smuggling step-42 alongside
off_allowlist_action run_shell_command is not on the allowlist
out_of_range_parameter target_lufs=-60.0 below min -31.0

And the guard fires on the real path, not just the test set. In a live run the agent submitted a conclusion blaming a preset it had not tested; the validator refused it, told it to call run_preset_experiment, and the agent corrected itself and resubmitted. That exchange is in the demo video.

The cost of agency, measured rather than estimated. The agent costs about seventeen seconds and seven Gemini calls more than a fixed sequence - and it decides its own path. RunConfig(max_llm_calls=14) is a hard bound; exhausting it ends the investigation with an escalation rather than a truncated answer.

Challenges I ran into

The ledger could not mint unique ids. Step ids were derived from the count of interpreted steps, so an agent running several tools before interpreting any of them minted step-01 every time and overwrote each observation's raw result.

The model blamed a preset it had not tested. So conclude now refuses an attribution with no experiment behind it. Then a later run showed the next failure mode: the model ignored that refusal five times in identical words and exhausted its budget. The second refusal now names the exact call to make instead of restating the rule.

Real footage exposed a defect a test pattern had hidden for the life of the project. Every normalize preset declared loudnorm=I=-23, and against a 440Hz sine that landed on -23.0 exactly. Single-pass loudnorm is a dynamic normaliser, though, and on real dialogue it delivered -20.4 LUFS - the normalise stage missing its own target by 2.6 LU against a +/-0.5 tolerance. resolve_audio_filter now measures first and applies a computed linear gain, which is what a broadcast normaliser does.

Cloud Run deployed a container with no media at all. gcloud run deploy --source . falls back to .gitignore when no .gcloudignore exists, and .gitignore excludes media/*.mp4. The fixtures are now generated at build time instead of shipped.

The evidence backend disappeared mid-project. The Grafana stack stopped answering between one day and the next, and three consecutive runs ended status: failed having already measured the file correctly and blocked it correctly. That is the wrong trade: the investigation is the only part that needs Grafana. It now degrades - it says the investigation did not happen, attributes no cause at all, and still proposes the repair the findings call for.

What I learned

Constraining an agent is not the opposite of making it agentic. The model here plans freely because it cannot execute anything, cannot decide compliance, and cannot cite evidence it did not cause to exist. Every removal of authority made it safe to give the model more room to think.

And schema validity proves shape, not entailment. The validator catches uncited, mis-cited and fabricated citations. It does not verify that the reasoning is sound - so the conclusion is deliberately hedged, and its weight comes from what the preset does, not from the model's confidence.

What's next

Channel-layout conformance, which would have caught this fault more cheaply than any investigation. A signed QC report artefact, since broadcast delivery is contractual. Severity levels - real QC has advisories, not just pass and fail.

Honest limits

  • Prepared scenarios are the default so results are reproducible. Uploading your own file runs the same pipeline against the default presets, with no fault injected - what you see is what the file really is
  • Three stages stand in for a nine-stage workflow (conform, colour, mix, mastering, versioning, transcode, wrap, package, deliver)
  • Correlation is not causation. A preset changing shortly before a failure is correlation; the conclusion is hedged accordingly
  • The measured +5.7 LU is a property of the test content, not of the preset: identical channels sum to +6.02 dB by arithmetic, decorrelated stereo to about +3 LU
  • The agent picks its own path, so two runs of the same scenario do not take the same route. That is the point, and it means a single run is not proof of a fixed behaviour

Note to judges

Reading is open, so the hosted page loads for anyone. Starting a run needs a demo token because every run costs real ffmpeg CPU and Vertex tokens - ask and I will share it. The Grafana dashboard link needs no login at all.

If the hosted demo is unavailable, the whole loop runs locally with no Grafana Cloud, no Vertex AI and no GCP project:

./scripts/make_fixtures.sh
docker compose -f docker/docker-compose.yml up -d
OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318 python scripts/demo.py --fixture fault

Narration in the video is a synthesised clone of my own voice. Programme footage is "Tears of Steel", (CC) Blender Foundation, mango.blender.org, used under CC-BY 3.0.

Built With

Share this project:

Updates

Submission history