Inspiration

"I must have gone through several hundred books on the subject, broken it down into categories on everything from his food tastes to the weather on the day of a specific battle, and cross-indexed all the data in a comprehensive research file." - Stanley Kubrick on researching Napoleon, to Joseph Gelmis, 1970

Stanley Kubrick kept about a thousand boxes of research per film. Clippings, photographs, books, interviews, maps. His team spent months building that archive before a single frame was shot. The archive was the edge.

Every filmmaker wants that edge. Almost none can afford it.

We wanted to know if an agent could build the archive overnight. A research department that keeps working unsupervised, notices what it still doesn't understand, and goes back out to fill the gap.

What it does

Type a film idea or premise and THE BOXES runs a loop on its own:

  1. Plans. Gemini 3.8 Flash writes a research ontology, one objective per box.
  2. Acquires. For each box, Parallel Search finds sources and Parallel Extract pulls the objective-specific full text. Every page gets mined for pictures, PDFs, audio, and video.
  3. Embeds. Gemini Embedding 2 turns every fragment into a vector. Text, images, PDFs, audio, and video land in one shared space. A reference photo sits next to an archival photo. A needle-drop sits next to a paragraph.
  4. Measures. Each objective gets a coverage score. A deterministic source-quality score rolls everything up into one research completeness number.
  5. Verifies. Embedding similarity finds pairs of evidence about the same thing. Gemini reads both and rules on whether they contradict, with an explanation and both citations.
  6. Opens its own boxes. When the evidence keeps circling a concept with no box, the agent proposes one and opens it. The loop stops when it's confident.

The Interface

The project page displays five core views above a persistent metric row:

  • Overview: A synthesized view of the world, highlighting the strongest evidence, thin boxes, and a grounded Ask interface that refrains from answering when data is lacking.
  • Departments: Expandable packets for each crew, containing box summaries, evidence, and inline media rendering.
  • Evidence: A filterable research map placing evidence fragments relative to their respective boxes. Contradictions carry clear tags and dedicated modal views. Users can manually upload references directly to the index.
  • Trace: A real-time console showing decision timelines, per-pass query ledgers, and verified contradiction verdicts.
  • Prior Art: A grid of existing films matching the premise by semantic similarity using TMDB and Gemini Embedding 2, paired with an analysis of unexplored creative angles.

How we built it

https://raw.githubusercontent.com/rapha18th/agentic-cinema-boxes/main/docs/architecture.png

THE BOXES architecture

The Boxes Agent: The production API runs the loop through a custom Google Agent Development Kit BaseAgent (ResearchWorkflowAgent) on Cloud Run. The same loop backs a conversational ADK tool surface: plan_research, run_research, query_index, and three more. Control flow is concurrent Python. Each round writes its search queries in one call, researches every objective in parallel, verifies every contradiction candidate in parallel, and embeds every fragment in parallel. Every step emits an event: plan, progress, search, extract, evidence, coverage, emergent_gap, contradiction, round_done, stop, reel.

  • Reasoning: Gemini 3.8 Flash handles objective planning, query drafting, contradiction analysis, and narrative generation via Vertex AI. Disabling extended thinking on structured outputs reduced call times from 20 seconds to 5 seconds.
  • Acquisition: Parallel Search and Extract run on every research round through the parallel-web SDK, with a raw HTTP path kept as a fallback. Search takes a semantic objective and returns ranked URLs with excerpts. Extract returns objective-specific full text, latency, result count, and extract status, which are recorded and shown in the trace. Parallel is the acquisition engine; the agent rewrites its Parallel objectives as its understanding develops. At kubrick depth, one Parallel Task API run answers a single focused question on the real history with its own citations, printed as a Deep Dive page in the dossier.
  • Multimodal embedding: Gemini Embedding 2 (gemini-embedding-2) produces 3,072 dimensions natively; the system stores the normalized 768-dimension cut in its own Firestore subcollection, kept off the browser-facing evidence documents. Parallel returns text. The system also harvests images (og:image, inline ), PDFs, and audio and video (, , og: audio, og: video, direct media links) from the pages Parallel surfaces. Text, images, and PDFs are harvested reliably from those pages. Audio and video mostly do not: the hosts that carry open recordings put the file behind a JavaScript player, and Parallel Extract returns markdown, so a page's element does not survive it. A direct API pass over Wikimedia Commons (File namespace, CC or public domain by policy) and archive.org (gated to explicit open licences and a short list of open collections) closes that gap: each objective tops up its audio and video budget from catalogues that hand back a fetchable file and a licence. ffmpeg trims the clip before embedding, with a container-aware byte trim as the fallback. Gemini Embedding 2 takes 8,192 tokens shared across modalities. The full asset stays linked at its source. Each non-text item embeds as media plus a caption.
  • Grounded Synthesis: The Ask interface uses cosine distance to retrieve top fragments. Gemini constructs answers strictly from retrieved context, providing sentence-level citations or abstaining when data is insufficient.
  • Metrics and Termination: Completeness scores factor in objective coverage, source tiering, source diversity, and contradiction penalties. The agent terminates execution when criteria cross defined thresholds.
  • Execution Stability: Distributed runs secure transactional Firestore leases to prevent race conditions. Interrupted SSE streams reattach to active background workers, while worker timeouts release stale leases automatically.
  • Prior Art Indexing: Candidate pools originate from TMDB and Parallel Search. Gemini Embedding 2 measures direct semantic proximity to the input premise, allowing Gemini to identify unclaimed narrative territory relative to existing works.
  • Document Generation: A background job compiles research dossiers into PDFs using reportlab, pairing synthesized text with visual moodboards for each research box.
  • Frontend and Storage: Built with React, Vite, and Tailwind CSS on Firebase Hosting. Authentication isolates user project data in Cloud Firestore and Cloud Storage, while execution logic runs on FastAPI inside Cloud Run.

Data sources

The live open web, reached through the Parallel Search and Extract APIs: newspapers, encyclopaedias, museum and national archive collections, oral history projects, government records, academic articles, and the images, PDFs, audio and video recordings those pages carry. TMDB for the prior-art candidate pool. User-uploaded references index alongside.

Challenges We Ran Into

  • Instruction-Based Embeddings: Gemini Embedding 2 lacks explicit task type parameters. Mapping text, images, documents, and trimmed media into a unified vector space required precise prompt prefixes and standardized asset preprocessing.
  • Media Processing Bottlenecks: Harvesting heavy media files slowed overall execution. Moving asset acquisition to dedicated I/O worker pools kept execution times manageable.
  • Client Thread Safety: Reusing a single genai.Client() across Python worker threads caused runtime failures. Resolving this required thread-local storage (threading.local()), enabling safe concurrent execution across search, verification, and embedding operations.
  • Rate Limits vs. Concurrency: Unbounded threading triggered Vertex AI rate limits, leading to execution delays caused by backoff sleeps. Bounding active network calls with a semaphore (threading.Semaphore(3)) reduced total execution time from 702 seconds to 347 seconds while processing a higher volume of assets.
  • Scraping Media Assets: Many institutional archives block direct media downloads or nest audio and video inside custom media players. Customizing user-agent headers and accept flags improved retrieval rates, though complete coverage requires headless browser processing.

Accomplishments that we're proud of

  • It opens its own boxes. A live run on a 1929 Vienna heist premise opened "VIENNA STOCK EXCHANGE AND SPECULATIVE PANIC" on its own. A radio-propaganda premise opened "AUDIO MATERIALITY AND TAPE ARCHIVES." Truly autonomous pre-production.
  • Unified Multimodal Indexing: Single runs successfully processed text, image, PDF, and video assets into a shared 768-dimensional space, referencing all modalities directly inside the generated summaries.
  • Automated Fact Verification: The system accurately identified and flagged direct historical contradictions between sources, such as conflicting construction dates for municipal housing projects.
  • Complete Data Provenance: Every fragment retains direct metadata linkages, including source URLs, dates, search queries, media types, and original asset locations.
  • Production-Ready Implementation: Delivered a fully deployed system featuring real-time event streaming, role-based access control, interactive visualization maps, and automated PDF dossier export.

What we learned

  • Memory, retrieval, routing, and judgement collapse into one operation on this platform: a distance in one vector space. Making that space multimodal changed what the tool can do.
  • Coverage and confidence numbers only earn their place if they're computed from something real. Tie them to an explicit plan and to verdicts.
  • Unbounded concurrency is a slower failure mode with extra steps. We timed it both ways.

What's next for The Boxes

  • Rights-aware reels. Track licence information per asset and assemble reels that respect it.
  • Vertex AI Vector Search for the multi-user index, and Agent Engine memory for cross-session recall.

Built With

Share this project:

Updates

Submission history