Inspiration

Physical film production runs on paper stripboards and gut instinct. When a lead actor gets a press hold, a permit falls through, or the weather turns, the 1st AD reshuffles days by hand — and there's no system of record for why a scene moved, whether the new plan is actually feasible, or what evidence justified the change. We wanted to see what an agentic system could do here without falling into the trap almost every "AI does X" demo falls into: letting a language model quietly invent facts it has no business inventing — a permit date, a day rate, a cast hold that doesn't exist.

What it does

SceneShift AI plans a tentative shooting schedule from real scenes, cast, locations, crew, and equipment; lets a producer ask What-If questions about disruptions in plain English ("What if Maya is out Saturday and the marsh floods Wednesday?"); and recovers three ranked alternative schedules, recommends one, waits for a named human's sign-off, executes it as a new schedule version, and verifies the result — cost delta, days delta, exactly which scenes moved. Every step writes an audit event.

Five scripted What-If scenarios exercise different disruption shapes, including one — an equipment failure with no public record — where the system deliberately says "no evidence found" instead of fabricating a claim. That's the detail we think matters most: most agent demos never show the honest failure case on purpose.

How we built it

The core architectural decision is a hard split: scheduling math is never delegated to a language model. A deterministic Python engine (constraint checking, packing, cost, scoring, diffs) owns every number. Gemini agents, via google-genai, only interpret producer language into a structured disruption spec, rank engine-generated recovery options, and explain tradeoffs in prose — they read production facts through typed tools and cannot write a scene, a cast hold, or a schedule block directly.

Real-world grounding for disruption claims — weather risk, permit windows, press holds — comes from Parallel's Search API, called live at runtime through an EvidenceProvider port. That port swaps between a live ParallelEvidenceProvider and a curated StubEvidenceProvider behind one factory function, so the full demo loop works identically offline, in CI, and in production.

Stack: FastAPI + SQLAlchemy + SQLite backend, React 19 + TypeScript + Vite frontend, deployed on Google Cloud Run (containerized FastAPI service) with Firebase Hosting serving the static frontend and rewriting /api/** straight through to Cloud Run — no separate API base URL to configure client-side.

Challenges we ran into

Getting a from-scratch full-stack app through an actual live Cloud Run + Firebase Hosting deploy — API enablement, Artifact Registry, Cloud Build, environment wiring for two separate live partner API keys — surfaces a lot of small friction points that never show up in a localhost demo: CORS origins that have to match the Hosting domain exactly, gcloud/firebase CLI auth quirks on Windows, and making sure the Hosting rewrite for /api/** doesn't accidentally also try to catch FastAPI's /docs route (it doesn't — that one has to be reached directly on the Cloud Run URL).

We also found, honestly, that our deterministic scoring engine's penalty weights are aggressive on a larger 40-scene board — Schedule Health and Resilience render 0/100 on the full Harbor Lights seed even on a fully feasible, zero-hard-violation schedule. Rather than hide it, we're calling it out here: it's a real, interesting tuning characteristic of an engine that refuses to let a language model paper over a low score with reassuring prose.

Accomplishments that we're proud of

A full product loop — Plan → Simulate → Recover → Approve → Execute → Verify → Audit — not a single proof-of-concept screen; a live production deployment with both Gemini and Parallel running for real, not fallback/stub, verifiable at /api/health; and a demo that's honest about a system limitation instead of cutting away from it.

What we learned

That the most convincing thing an agentic system can do is refuse to guess. The "no evidence found" scenario got more attention in our own internal reviews than any of the four scenarios where the system found something.

What's next for SceneShift AI

Grafana Cloud MCP as a live observability sink for agent traces (sketched, not yet implemented), richer stripboard editing and multi-production compare, and a tuning pass on the scoring engine's penalty weights for larger boards.

Built With

Share this project:

Updates

Submission history