Inspiration
It started with a scenario that's almost boringly common on a film set: a location falls through. Not a dramatic disaster — just a permit issue, or the tide schedule changes, or the beach you booked is suddenly closed for the week. Small problem. Except it isn't small, because that one change has to ripple through budget, scheduling, storyboards, music, risk assessment — and right now, that ripple happens by hand. Ten to twenty people re-reading the same screenplay, each reconstructing their own piece of it, hoping nothing falls through the cracks in the group chat.
We wanted to know: what if the screenplay only had to be understood once? Not summarized once — understood, structurally, by a system that knows which departments actually care about which changes, and updates only those.
What it does
CinemaPilot is a production office made of 14 collaborating agents, all reading from and writing to one shared Production Graph in BigQuery. A Script Intelligence agent extracts a screenplay into structured data — scenes, characters, locations, props, timeline — once. From there, everything is event-sourced: every change is logged, and a Change Detection agent diffs the graph and computes exactly which downstream agents actually need to react.
Move Scene 5 from a storage unit to a beach, and you don't get all 14 agents lighting up — you get 6: Budget recalculates cost against real location data, Risk flags real logistics concerns (tide windows, permits, no power on-site) and escalates the serious ones straight into Grafana as real Incidents, Schedule reworks the shoot day, Storyboard generates a new grounded concept panel, and Music composes a fresh cue for the new setting. A Producer agent synthesizes it all into an overview, and — this is the part we're most proud of — a Producer readiness gate actually queries Grafana live and refuses to greenlight the shoot while a safety incident is still open. Resolve the incident, and the gate flips to ready in real time.
There's also a Trailer agent that assembles a concept trailer from the storyboards using Veo 3.1 for motion and Lyria for score — launched on-demand from the dashboard, with an honest, visible fallback to a local ffmpeg animatic when Veo access isn't configured, rather than pretending.
Built for filmmakers, screenwriters, and studio production crews — anyone who's felt a single script change turn into a day of manual cross-department scrambling.
How we built it
Gemini does the reasoning across every agent. Google ADK handles tool-calling into the Grafana Cloud MCP server — that's not a side integration, it's load-bearing: the Risk agent creates and deduplicates real Grafana Incidents, the Location agent queries live Prometheus and Loki data, and the Producer's readiness gate is a genuine runtime dependency on Grafana being reachable. BigQuery holds the Production Graph as an append-only event log, which is what let us build a Change Detection agent that diffs cleanly instead of guessing. Storyboards come from Gemini's native image generation, music from Lyria 3, video from Veo 3.1, and the whole dashboard runs on Cloud Run, with secrets in Secret Manager and IAM scoped down to what each service account actually needs.
Challenges we ran into
Imagen 3 was the first wall — it's been quietly deprecated by Google and returns a flat 404 on Vertex AI for any new project, which we didn't discover until our Storyboard agent had been failing mid-build. We switched to Gemini's native image generation and disclosed the swap in the README rather than pretend the original spec still stood.
That failure taught us something we then applied everywhere: "the code ran without an exception" and "the output is correct" are different claims, and conflating them is exactly how a broken agent ships unnoticed. From then on, every agent got the same treatment — real byte-level checks on generated assets, real dashboard rendering confirmed, real Grafana incidents queried directly, not inferred from logs.
Determinism was its own puzzle. Even at temperature=0.0, our Budget agent occasionally flipped between two cost estimates, because the model was emitting its number before any reasoning tokens — a near-tied-logit problem, not something a retry would fix. Reordering the schema so reasoning comes first, and replacing two competing cost heuristics with one explicit formula, got us to 10/10 reproducible runs.
Cost discipline mattered too, especially for Veo. Real video generation runs several dollars per trailer, so the Trailer agent is deliberately on-demand only — never part of the automatic cascade — and falls back visibly to a local ffmpeg animatic when Veo isn't configured, instead of failing silently or faking success.
Accomplishments we're proud of
The Producer agent's readiness gate is the one we're most proud of: it queries Grafana Cloud's Incidents and Prometheus data live, over MCP, and will not greenlight a shoot while a real safety incident is open. That's Grafana doing actual work in the system, not a dashboard bolted on to satisfy a track requirement. We're also proud of catching our own mistakes before anyone else had to — the Imagen deprecation, a duplicate-incident bug in the Risk agent, and the Budget agent's non-determinism were all found and fixed through our own verification, not left for a judge to discover.
What we learned
That trust has to be earned from evidence, not assumed from a clean exit code — a lesson expensive enough that it reshaped how we built the remaining agents. We also learned that "deterministic" is a much narrower promise than temperature=0.0 implies once a model is choosing between two plausible answers, and that fixing it means removing the ambiguity from the prompt, not just constraining the sampling. And practically: read a model's deprecation notices before architecting three agents around it.
What's next
Real-time collaborative editing of the Production Graph across multiple producers working the same production at once, and extending Change Detection's routing rules beyond location and character changes to cover the fuller structure of a screenplay — dialogue changes, prop substitutions, timeline reordering — so more of a production's actual day-to-day disruptions get the same automatic, minimal-blast-radius treatment Scene 5 gets today.
Built With
- bigquery
- cloud-run
- document-ai
- fastapi
- ffmpeg
- gemini
- google-adk
- google-cloud
- google-cloud-aiplatform
- google-cloud-secret-manager
- google-genai
- grafana-cloud-mcp
- lyria
- opentelemetry
- python
- veo

Log in or sign up for Devpost to join the conversation.