Inspiration
You can hear every line of a film and still miss the moment someone decides to leave. Audio description is the narration track that carries those moments, and almost nobody ships it, because it is expensive. A writer watches the film, finds every pause in the dialogue, and writes copy that fits that pause exactly. Too long and it talks over the next line. Too vague and the viewer loses the plot.
That constraint is the whole problem: tight, measurable, unforgiving. Not "summarise this video," but "you have 4.0 seconds and 11 words, what matters most?" We wanted to know whether an agent crew could hold that honestly, instead of generating plausible narration and hoping it fit.
What it does
EarSight takes a video and returns a described version: the original audio with narration mixed into the silences, plus a WebVTT description track.
The pipeline is silence-first, and that is the design decision everything else follows from. It transcribes the dialogue with word-level timestamps using Gemini on Vertex AI, finds every gap of 0.5 seconds or more with no spoken word in it, sets a hard word budget for each gap of floor(seconds x 2.75), watches what happens on screen inside that gap, writes copy to the budget, speaks it with Gemini 3.1 Flash TTS in a calm documentary register, then ducks the original audio and mixes with ffmpeg.
The budget is a ceiling, not a suggestion. Copy that exceeds it is dropped, never truncated, because a sentence that stops mid-word is worse than silence. The model must call check_budget and check_collision as forced function tools before it can emit a cue, and the same checks run again as Python assertions before anything is published. The model can be persuasive. The assertion cannot.
How we built it
Six Python services, one per stage, each its own Cloud Run container: orchestrator, transcriber, framer, describer, synthesiser, mixer. Every handoff crosses a Confluent Kafka topic, eight in total including a dead-letter. Nothing is mocked; the topics are load-bearing.
The describer is the only agent that reasons. It is a native Google ADK LlmAgent with the budget and collision guards exposed as function tools under tool_config mode=ANY, so the safety checks cannot be skipped. It watches the source video straight from Cloud Storage through Gemini's native video understanding rather than a still frame, because a still cannot tell a hand reaching from a hand withdrawing, and that distinction is most of what audio description exists to convey.
The frontend is Next.js on Cloud Run: a time-proportional waveform showing dialogue regions and silence windows filling as cues are placed, a live lane of messages crossing the topics, and a custom keyboard-accessible player.
The whole project was built with IBM Bob, one step per session against a written plan, with every step's reasoning committed alongside the code.
Challenges we ran into
Almost everything that broke, broke only in production. The pipeline ran locally for days. The moment we deployed it, six distinct bugs surfaced that no local run could ever have caught, and every one of them was ours.
We passed the Cloud Run region straight through as the Vertex AI location, so our Gemini 3.x calls never reached the endpoint we intended. We built a Vertex client with workload identity and then let the agent construct its own, which is how we ended up staring at "No API key was provided" in an environment where we had deliberately removed the key. Our waveform never rendered because of one wrong ffmpeg flag, and the error was swallowed as non-fatal, so the peaks file was never written and nobody noticed. None of our Docker images set PYTHONUNBUFFERED, so every log line in every service sat in a buffer instead of reaching Cloud Logging, and we spent a day debugging a distributed system blind. Our deploy config was invalid YAML and had never once been executed.
And the one that mattered most: the describer takes minutes of model inference per job, and we were doing that work inside the Kafka consumer poll loop, past the default five-minute interval. The broker did exactly the right thing and evicted a consumer that had stopped responding. Our dead-letter topic caught the failure with its reason attached, which is how we found it.
Bob is what made those six days survivable. Every fix landed as its own step against the written plan, with the reasoning committed next to the code, so we could always see what we had already ruled out instead of rediscovering it.
What we learned
Deploy earlier. Every serious bug we hit lived in the seam between services, in the code path that only executes in the deployed system. A local script exercised our logic modules perfectly and told us nothing about whether the product worked.
Constrain the model, then verify it anyway. Forced function calling on Gemini makes the model consult the rules before it writes; a deterministic assertion is what makes the rules true. Both, not either. That pairing is the reason we can promise a studio that no cue will ever talk over a line of dialogue.
Use each primitive for what it is good at. Kafka gave us durability, backpressure and a dead-letter queue that earned its keep on day one. Long inference belongs beside the consumer, not inside it, and knowing that now is what lets the same architecture carry a whole catalogue rather than a single file.
And build with a record. Working with Bob one step per session meant the project arrived with its own history: the plan written before any code, seven reasoning documents, five encoded workflow rules, a session log with 31 entries, and 21 of 22 commits signed with how the decision was made.
What's next for EarSight
Extended description. Real audio description sometimes needs more room than the film leaves, and the standard answer is to hold the picture for a beat. Because we already know the exact budget of every silence, we know precisely where a scene needs that beat and how long it should be, so the viewer gets the whole story rather than the part that fit.
Description in any language. The silences do not move when the language does. Gemini already writes and speaks multilingually, so one described master can ship as a Spanish, Arabic or Japanese track from the same timings, which turns accessibility work into international distribution.
Voice as part of the house style. Per-title voice and register so a nature documentary and a thriller do not sound like the same narrator, and a second voice reserved for on-screen text and signage so viewers can hear the difference between the world and the words printed on it.
An editor, not just an output. Professional describers are the people who should have the last word, so the next release hands them our pass as a first draft in a timeline where the word budget is enforced live as they type. Hours of writing become minutes of review.
And the whole back catalogue. Thousands of titles that have never been described, processed in parallel and delivered as WebVTT any existing player can already read. That is the point where the event-driven architecture stops being a design choice and starts being the reason the work is possible at all.
Built With
- apache-kafka
- cloud-build
- cloud-logging
- cloud-run
- cloud-storage
- confluent
- docker
- fastapi
- ffmpeg
- gemini
- gemini-tts
- github-actions
- google-adk
- google-cloud
- ibm-bob
- nextjs
- playwright
- python
- react
- secret-manager
- typescript
- vertex-ai
- video-understanding
- webvtt


Log in or sign up for Devpost to join the conversation.