Inspiration
Film pre-production is notoriously chaotic. Hundreds of creative and technical decisions, including scene breakdowns, location permits, camera rigs, stunt safety protocols, and VFX plate requirements, are fragmented across scattered emails, PDFs, Final Draft scripts, Excel budgets, Word treatments, and messaging channels. When production schedules slip or safety hazards are overlooked, budgets bleed and set safety is jeopardized.
While generative AI models have demonstrated extraordinary creative capabilities in scriptwriting and concept generation, Hollywood and professional crews face a critical barrier to adoption: AI hallucination, lack of operational transparency, and the absence of verifiable lineage. Production heads cannot risk a crew's safety or a multi-million-dollar shooting day on an untraceable, black-box AI output.
We asked ourselves: What if creative crews had a unified, real-time collaboration space (familiar as a modern chat app) where native Gemini AI agents collaborate alongside human crew leaders, backed by deterministic document intelligence across all industry formats, immutable artifact lineage, and enterprise-grade observability from Grafana Cloud?
That vision is StudioTower.
What it does
StudioTower is an observable, agentic control tower for film preparation and production coordination that bridges creative crew collaboration and industrial-grade reliability:
Chat-Centric Crew Collaboration Spaces:
- Organized into secure, multi-tenant Spaces with flexible team permissions, designed with the intuitive ease of Slack, Discord, Google Chat, or LINE groups.
- Intra-Space Project & Topic Tagging: Teams can isolate and pivot conversations by project blocks, shoot days, or departments (e.g.,
#block-a,#stunts,#vfx,#general) without fragmenting the crew across siloed channels.
Native AI Crew Agents in the Conversation:
- Team members can directly dispatch complex production tasks to an AI agent right in the chat stream.
- Powered by Gemini 3.6 (
gemini-3.6-flash) through thegoogle-genaiSDK. StudioTower embodies a virtual crew leadership council as one structured system prompt (not five separately scheduled ADK agents):- Production Director: Guarding artistic vision and character beats.
- 1st AD & Line Producer: Calculating logistical feasibility and day-out-of-days complexity.
- Director of Photography (DP) & Key Grip: Outlining camera setups, lighting, and unit moves.
- Stunt & Safety Coordinator: Flagging hazardous stunts, pyro, and location risks.
- VFX Supervisor: Categorizing plate requirements and greenscreen tiers.
Multi-Format Document Intake & Asset Generation via ProDocuX:
- Deep Multi-Format Ingestion: StudioTower calls PyPI
prodocux==0.3.0rc5(prodocux_kernel) for the formats Kernel owns. Locators stay citation-stable (page:N,sheet:{name}:rows:{N}-{M},slide:N). Formula-like cells are flagged, not rewritten.- PDF: Kernel
extract_pdf_byteswithpage:Nlocators. - DOCX: Kernel
extract_content_blocksinto section/table chunks. - CSV & XLSX: Kernel
extract_content_blocks(rows:N-M/sheet:{name}:rows:{N}-{M}). - PPTX: Kernel
profile_pptx_bytes(slide:N) - text, tables, speaker notes, image counts; not a vision model. Kernelextract_content_blocksTypeError's on PPTX tables in 0.3.0rc5. - FDX (Final Draft) and TXT: no Kernel extractor; StudioTower parsers keep
scene:N:SLUGLINE/section:N.
- PDF: Kernel
- Automated Deliverable Generation: StudioTower builds evidence-derived content blocks and ProDocuX Kernel writes the user-selected format (
pdf,docx,xlsx,pptx,csv, orjson). These are source-grounded working files. Kernel writers exist for those formats; crew-ready department templates (call sheets, breakdowns) still need to be prepared against real production layouts. - PDF, Word, spreadsheet, deck, and CSV outputs share that Kernel writer path. JSON manifests remain for workflow control.
- Deep Multi-Format Ingestion: StudioTower calls PyPI
Cryptographic Lineage & Human Approval Gates:
- User-facing deliverables are rendered by ProDocuX Kernel. After a human approval gate, PDX Artifact Engine (
pdx-artifact-engine==0.3.0a6) validates apdx_execution_plan_v1andArtifactRuntime.execute_planwrites the control package: Shoot Schedule CSV, Conflict Matrix JSON, Unit Handoff Manifest JSON, and a checksummed run manifest. The call sheet / breakdown / budget file itself is not re-executed inside PDX Runtime. - High-risk operations (e.g., live fire stunts, extreme tech asset locks) trigger mandatory Risk Gates & Human Approvals, preventing unverified autonomous actions.
- Hackathon inspector Lineage tab: Grafana Cloud Tempo spans for the active Space (
studiotower.space.eventwithspace_id), rendered as a process flow (file → message → run → gate) from Grafana HTTP API trace attributes. StudioTower still stores a local artifact DAG (source file ➔ chunks ➔ run ➔ PDX manifests); that graph is parked in the hackathon UI so the partner-track surface is Grafana-only.
- User-facing deliverables are rendered by ProDocuX Kernel. After a human approval gate, PDX Artifact Engine (
Observability & Agentic Diagnosis via Grafana Cloud:
- Production Cloud Run runs the official
grafana/mcp-grafanaserver (Streamable HTTP) next to FastAPI, authenticated to Grafana Cloud stackhttps://loftyladybug3305.grafana.netwith a Grafana service account token (glsa_) for reads. - Space activity is exported over Grafana Cloud OTLP (
glc_write token) into Prometheus and Tempo. Cloud Run flushes spans and metrics on the request path so traces survivemin-instances=0. - The right-panel Telemetry tab loads Prometheus Space-activity stats and a timeseries through the Grafana HTTP API (
/api/search,/api/dashboards/uid,/api/ds/query). Charts are Grafana query frames. Empty panels mean Grafana returned no frames, not invented data. The inspector does not print the Grafana stack hostname. StudioTower run cards and the local span waterfall are parked in this hackathon build. - On diagnosis, StudioTower also issues MCP JSON-RPC
initialize,tools/list, andtools/call. Hostedmcp.grafana.combrowser OAuth is not used for the unattended API.
- Production Cloud Run runs the official
How we built it
StudioTower was built from the ground up as a cloud-native, clean-room architecture:
- AI Reasoning: Google Gemini 3.6 (
gemini-3.6-flash) viagoogle-genai, with Pydantic structured output schemas (SceneBreakdown,ResourceRequirement,ConflictItem,RiskGateProposal) and a single multi-persona crew prompt. - Observability Partner Track (Grafana Labs):
- Official
grafana/mcp-grafanasidecar on Cloud Run (127.0.0.1:8000/mcp). - Grafana HTTP API boards in the hackathon Lineage tab (Tempo process flow) and Telemetry tab (Prometheus stats + trend) via
/api/searchand/api/ds/query. - Grafana Cloud OTLP write (
glc_) for Space activity metrics and traces; gateway region is taken from the stack Prometheus URL (this demo:prod-us-east-2). - Diagnosis still attaches MCP
tools/callevidence when the sidecar andglsa_token are configured. Local span waterfall is parked for the hackathon UI.
- Official
- Deterministic Pipeline & Asset Engines:
- ProDocuX Kernel: PyPI
prodocux==0.3.0rc5extract/profile APIs for PDF, DOCX, CSV, XLSX, and PPTX. FDX/TXT remain StudioTower because Kernel has no APIs for those formats. - PDX Artifact Engine: PyPI
pdx-artifact-engine==0.3.0a6(validate_execution_plan,ArtifactRuntime.execute_plan) plus StudioTower human approval gates.
- ProDocuX Kernel: PyPI
- Backend & Cloud Infrastructure:
- Python 3.12 & FastAPI hosted on Google Cloud Run.
- Google Cloud Firestore for durable multi-tenant space, message, run, and membership state.
- Google Cloud Storage (GCS) for uploaded production scripts, media assets, and output manifests (SHA-256 lineage; not object-lock / WORM).
- Firebase Authentication providing multi-user identity and role-based access control.
- Frontend SPA: TypeScript + Preact / Vite interface featuring a high-contrast dark theme, live Markdown streaming, responsive chat composer, and Grafana Cloud boards in the inspector. HITL approvals live in the header Tasks & Approvals view.
Challenges we ran into
- Bridging Non-Deterministic AI with Deterministic Multi-Format Document Realities:
- Film productions rely on heterogeneous document types, from Final Draft XML to Excel spreadsheets and scanned PDF callsheets. Gemini 3.6 is grounded on Kernel extract/profile locators for PDF/DOCX/CSV/XLSX/PPTX; Final Draft remains a StudioTower parser because Kernel 0.3.0rc5 has no FDX API.
- Real-Time Telemetry Correlation in Asynchronous Agent Loops:
- Agentic multi-turn workflows generate nested calls (document parsing, multi-role reasoning, risk validation, file persistence). Every child span must keep causality under a unified
trace_idandrun_id. Hosted Grafana Cloud MCP is OAuth-browser-only, so unattended Cloud Run uses the officialgrafana/mcp-grafanasidecar with a service-account token, and diagnosis fail-closes when MCP is missing rather than inventing Grafana data.
- Agentic multi-turn workflows generate nested calls (document parsing, multi-role reasoning, risk validation, file persistence). Every child span must keep causality under a unified
- Multi-Tenant Security with Flexible Crew Access:
- Film productions constantly bring in external contractors (stunt teams, prop masters, VFX vendors) who should only access specific topics or files. StudioTower enforces binary Space membership plus intra-space tag filtering (
#stunts,#vfx) and signed token verification. It does not yet implement per-file or per-topic contractor ACLs inside a Space.
- Film productions constantly bring in external contractors (stunt teams, prop masters, VFX vendors) who should only access specific topics or files. StudioTower enforces binary Space membership plus intra-space tag filtering (
- Making Grafana Cloud the live inspector, not a deep link:
- Grafana read (
glsa_HTTP API / MCP) and write (OTLPglc_) are different credentials. A read-only service account cannot ingest traces. - OTLP Basic-auth username is the Grafana Cloud instance id, not the org id, and not always the Prometheus datasource
basicAuthUser. - A hardcoded OTLP gateway region (
prod-us-east-0) disagreed with this stack's Prometheus/Tempo hosts (prod-us-east-2); the wrong gateway authenticated some tenants and still dropped data. - Cloud Run allocates CPU only during the request.
BatchSpanProcessordelayed export until after freeze, so Tempo stayed empty until request-path flush.
- Grafana read (
Accomplishments that we're proud of
- True Enterprise Observability in Agentic Cinema: Official grafana/mcp-grafana at runtime, diagnosis via MCP
tools/call, and in-app Grafana HTTP API boards that show this Space's Prometheus events and Tempo traces after OTLP ingest - not Billing/Usage dashboards, not invented series. - ProDocuX + PDX Artifact Engine on the live path: Kernel extract/profile and writers for PDF/Office/CSV, plus
ArtifactRuntime.execute_planfor checksummed control manifests after human approval. - Authenticated human approval gates: No stunt or safety action can execute without an authenticated human approval recorded with actor identity and artifact lineage.
- Source-grounded working files: Incoming PDF/FDX/DOCX/XLSX/PPTX parsed with stable locators; Kernel writes the selected format as a working file rather than a black-box PDF.
What we learned
- Observability is the missing link for generative AI adoption: Film producers are open to AI, but only when they can inspect execution traces, verify data sources, and monitor latency and costs via dashboards like Grafana.
- Document ground-truth requires precise locators: Storing chunk offsets as explicit
page:N,scene:N, orrows:N-Mlocators allows AI agents to cite exact script evidence, eliminating crew skepticism. - Human-in-the-loop is not a limitation, it's a feature: Creative teams demand agency. Empowering crews to approve or reject proposed AI action cards creates trust and accelerates real-world workflow integration.
What's next for StudioTower: Observable Agentic Cinema Control
We envision StudioTower expanding along two powerful development trajectories.
Near-term StudioTower engineering (supports those trajectories; does not replace them): keep Grafana Cloud OTLP aligned to the stack Prometheus region with request-path flush on Cloud Run; optionally restore the local artifact DAG and span waterfall beside the Grafana boards; layer crew-department document templates on the existing ProDocuX Kernel writers.
1. Vertical Specialization: Deep Agentic Cinema Ecosystem
- Generative Media Pipeline Integration: Connect StudioTower's approved breakdowns directly to Google Cloud generative media models (e.g., Imagen, Veo) and AI music/scoring tools to generate instant animatics, concept art, and temp soundtracks directly from approved scene nodes.
- Automated Call Sheet & Scheduling Matrix: Dynamically synthesize crew schedules, golden-hour sun positions, weather forecasts, and talent union turnaround rules into automated, downloadable industry-standard Call Sheets.
- Smart Set Telemetry: Ingest real-time camera metadata, sound logs, and DIT offload statuses into Grafana Cloud dashboards to monitor active shooting days live.
2. Horizontal Generalization: Universal Observable Agentic Workspaces
- Enterprise Control Tower Platform: Generalize the core underlying stack, chat-centric team spaces, multi-format document intake (PDF, Office, CSV), cryptographic artifact DAGs, human approval gates, and Grafana MCP observability, into a universal platform for other mission-critical industries.
- Cross-Industry Applications: Adapt StudioTower's architecture to Legal Discovery (case briefs and evidence lineage), Emergency Response Coordination (hazard assessment and multi-agency gates), and Architecture & Engineering Construction (AEC blueprint approvals and safety audits).
Built With
- cloud-tasks
- fastapi
- firebase
- firestore
- gemini
- google-cloud
- google-cloud-run
- grafana
- mcp
- model-context-protocol
- opentelemetry
- pdx-artifact-engine
- preact
- prodocux
- pydantic
- python
- typescript
- vite
Log in or sign up for Devpost to join the conversation.