Inspiration
Film and TV production is expensive. A single continuity error an actor's coffee mug switching hands between takes, a jacket that appears and disappears, a prop that moves across the room can mean an entire scene needs to be reshot. Script supervisors catch these on set by hand, reviewing notes, photos, and memory against what they see in the monitor. It's painstaking, high-stakes work done by humans under time pressure. We wanted to explore what happens when you give that role to a multimodal AI. Gemini can see video. It can read a screenplay. It can compare two things and articulate exactly what changed. The question was: could we build a tool that bridges those capabilities into something a real production could use in under a day of development?
What it does
Continuity Guardian is a three-step AI pipeline for catching continuity errors between video takes:
Parse Script — upload a
.txtor.pdfscreenplay; a Gemini agent extracts every scene's slug line, characters, props, costume notes, and action summary into structured JSON.Analyze Takes — upload two video clips (a Reference take and a Check take) for the same scene; a second Gemini agent visually inspects each clip and produces a detailed
TakeDescriptioncovering props in frame, costume observations, blocking, framing, and a free-form narrative.Generate Report — a third Gemini agent compares the two
TakeDescriptionobjects field-by-field and produces aContinuityReportwith every discrepancy categorised by type (props,costume,blocking,character,other), severity (low,medium,high), affected character, and plain-English description — plus an overall PASS / FAIL verdict.
The entire workflow is accessible through a web dashboard served by the same FastAPI application. No external database is required.
How we built it
Agent architecture
Each of the three pipeline stages is a native Google ADK LlmAgent with a Pydantic output_schema
Using output_schema means the framework enforces the output shape we never parse free-form model text in the business logic. Each agent runs in its own InMemoryRunner with a fresh session_id per request, so there is no state contamination between calls.
The structured output contract for the comparison step is: \text{ContinuityReport} = {\ \text{discrepancies},\ \text{overall_severity},\ \text{pass_fail},\ \text{summary}\ }
where each discrepancy carries: d_i = (\text{category},\ \text{severity},\ \text{description},\ \text{affected_character})
and the verdict is defined as: \text{pass_fail} = \begin{cases} \text{PASS} & \text{if } \nexists\ d_i : \text{severity}(d_i) = \text{high} \ \text{FAIL} & \text{otherwise} \end{cases}
IBM Bob
IBM Bob was used as the primary development partner throughout the build. Bob scaffolded the entire project from a single spec prompt generating all agent files, the FastAPI routes, the Pydantic schemas, the Confluent producer/consumer, the dashboard, Dockerfile, and .env.example in one session. Subsequent sessions with Bob resolved the ADK session state internals (reading from session.state[output_key] rather than parsing event.content.parts text), iterated on the dashboard UX, and produced this documentation.
Challenges we ran into
ADK output_schema and session state
The biggest technical challenge was understanding exactly how ADK surfaces output_schema results. When an LlmAgent has output_schema set, the validated result is written into session.state[output_key] not into event.content.parts[0].text as one might expect from the streaming API. Our initial implementation read the text part and tried to parse it as JSON, which worked sometimes but failed silently when the model streamed partial tokens or wrapped the output in markdown fences. The fix was to read from session state as the primary path, with text parsing only as a fallback.
Multimodal video payloads
Sending video bytes through FastAPI's multipart upload to the ADK InMemoryRunner required constructing genai_types.Part.from_bytes with the correct MIME type. Video files are large a 30-second .mp4 can be 50–100 MB and the entire payload passes through the Gemini API inline. This works for a hackathon but would need chunking or Cloud Storage URI references for production use.
Browser caching during iterative development
Static files served by FastAPI's StaticFiles were aggressively cached by the browser. During rapid dashboard iteration, the browser would serve old JS against new HTML, causing null reference errors that looked like code bugs. The fix was adding ?v=N cache-bust query strings to <script> and <link> tags, and using optional chaining (?.addEventListener) throughout the JS so a stale cache can never crash the entire script.
Accomplishments that we're proud of
- Three agents, one pipeline, zero manual wiring — the ADK
output_schemachain means the output of one agent feeds directly into the next as validated Python dicts, with no prompt engineering needed to coerce the model into a particular format. - A working multimodal video analysis pipeline — uploading a real video clip and getting back a structured, field-level visual description of its contents in under two minutes is genuinely impressive and was not obvious that it would work reliably.
- A fully self-contained deployment — one
Dockerfile, onegcloud run deploycommand, and the entire application (dashboard, API, agents) is live on a public URL with no infrastructure to manage. - Optional Kafka integration without breaking the default path —
USE_CONFLUENT=falsemeans the app runs identically whether or not Kafka is configured. The producer is a pure no-op.
What we learned
- ADK
InMemoryRunneris request-scoped by design — creating a fresh runner and session per API call is the correct pattern, not a workaround. It gives you complete isolation between concurrent requests with no shared state. -output_schemaenforces structure at the framework level — you do not need to ask the model to "return JSON" in the prompt. The schema constraint is applied by ADK before the result reaches your code. - Gemini 2.5 Flash handles video surprisingly well for structured tasks — given a precise instruction and a constrained output schema, it reliably identifies props, costume details, and blocking from short video clips. Longer clips and complex multi-character scenes are where accuracy starts to degrade.
- Separating the build from the runtime configuration — understanding that
docker buildpackages code andgcloud run deploy --set-env-varsinjects secrets at runtime (so the image never contains credentials) is a foundational cloud-native pattern worth internalising early.
What's next for continuity-guardian
-Firestore persistence — replace the JSON file store with Firestore so data survives redeployments and scales across multiple Cloud Run instances.
- Scene-level comparison — compare a take directly against the parsed script notes for that scene (
POST /reports/compare-to-scriptis already implemented in the API but not yet surfaced in the dashboard). - Thumbnail extraction — extract a key frame from each take and display it alongside the continuity report so the supervisor can see the discrepancy visually, not just read about it.
- Batch scene processing — process all takes for an entire shooting day in one job, producing a consolidated daily continuity log.
- Vertex AI Agent Engine deployment — deploy the
continuity_report_agentto Agent Engine for managed scaling, built-in tracing, and integration with the rest of the Google Cloud AI ecosystem. - Confidence scoring — extend the
Discrepancyschema with a model confidence field so supervisors know when to trust the AI's finding versus when to verify manually.
Built With
- bob
- confluent
- fastapi
- gemini
- google-adk
- google-cloud-run
- ibm
Log in or sign up for Devpost to join the conversation.