Inspiration

Every AI video tool gives you a clip. None of them give you a production.

A real studio doesn't start with a prompt. It starts with research — what does the audience actually want right now? Then it builds a world, keeps every character on-model across every shot, quality-checks the footage, assembles it, narrates it, and ships a campaign.

That whole loop is the bottleneck. We wanted to find out whether a network of agents could hold creative context from the first market signal all the way to the final post — not generate a clip, but run a studio.

What it does

You type one sentence. A network of agents then:

  1. Discovers — Parallel Search scans live web signals for a real content opportunity
  2. Thinks — Gemini turns those signals into a content angle and creative strategy
  3. Creates — builds a persistent World Bible: characters, locations, visual rules
  4. Produces — frames → Veo video → VLM continuity QC → FFmpeg assembly
  5. Distributes — an 8-platform campaign with on-model creatives and copy

Expensive steps sit behind human approval gates, and every asset is costed in a transparent per-asset ledger before a cent is spent.

Beyond the pipeline, 11 standalone tools work on their own — Text→Image, Text→Video, Image→Video, Voiceover+Images→Video, Text→Speech, Image→Prompt, Upscale, Social Post, YouTube Kit, Creative Text Editor, and Cast (build a character once, reuse it in any scene).

How we built it

Five planning agents hand each other schema-validated objects, not prose, so a later stage can rely on the shape of what it receives:

Agent Produces
research-agent Evidence gathered through Parallel Search, carrying its sources
opportunity-agent The angle worth making, argued from that evidence
world-builder-agent The World Bible — characters, locations, props, visual language
storyboard-agent Shots, with durations reconciled to the requested runtime
production-planner-agent Per-asset model and cost plan

Every model in the system is a Google model — Gemini on Vertex AI for reasoning, gemini-2.5-flash-image for stills, Veo 3.1 for video, Gemini VLM for continuity QC, Gemini TTS for narration. Firebase handles auth, Firestore and Storage. There are no third-party AI providers anywhere in the stack.

QC is a gate, not a report. Every generated shot goes back to a vision model holding the World Bible and is asked whether the character, location and look actually match. A pass moves forward; a fail regenerates that shot against the same reference, bounded so a stubborn shot can't burn the budget.

Every provider sits behind an interface chosen by one environment variable, so the entire pipeline runs end to end with zero API spend. That's how it was developed.

Challenges we ran into

The Parallel integration looked like it worked, and didn't. We were sending the planned queries as { query: "..." }. The API's field is search_queries — an array. Parallel ignores unknown fields rather than rejecting them, so every request returned HTTP 200 with plausible, well-formed results. They were just ranked against our hardcoded objective alone, identically for every run. We found it by searching something deliberately absurd: asking for "competitive axe throwing league Estonia" returned articles about social media trends. The same key sending search_queries returned axe-throwing venues in Tallinn. A second bug hid under it — result bodies arrive in excerpts, an array, and we were reading snippet, so every piece of evidence carried an empty body and the synthesis step was reasoning over bare titles while appearing to cite sources correctly. Both were silent. Neither would ever have thrown.

Fixing the research broke the pipeline. With real excerpts in the prompt, synthesis went from 6 seconds to 36. The six-call planning chain stopped fitting inside a serverless function's execution limit, and the function was terminated rather than throwing — so nothing was ever written as an error, and the UI waited forever on a status that could never change. Every stage already persisted its own output, so we made each one skip itself when its output exists and added a resume path the client triggers when a run stops advancing. Completed work is never re-run, so a resume can't duplicate spend.

One credential resolver served two clouds. When we moved Vertex to a separate Google Cloud project from Firebase, the shared service-account resolver preferred the Vertex credential — and it was also used by Firebase Auth, Firestore and Storage. The project id still came from the Firebase env var, so nothing looked misrouted; sign-in, credits and every asset write would simply have started failing against a project that account had no permissions in. Model plane and data plane are now separate resolvers with opposite precedence.

firebase-admin reaches jose through a CommonJS require. On Node below 22.12 that throws ERR_REQUIRE_ESM — and the app otherwise appears to start normally while every route touching Auth or Firestore returns a blank 500 with a minified React error. We pinned the Node floor and added a health endpoint that actually initialises Firebase Admin rather than just parsing the key.

Smaller ones that each cost a day: signed URLs expire, so a 7-day URL persisted into a database is a 404 with a delay fuse — assets now resolve through a route. Data URLs above a few hundred KB blow the server-action argument limit, which silently broke every image-taking tool for real photographs. The cost ledger is read-modify-write, so a retry mid-run clobbered the run's own accounting until we serialised per-project writes. And an LLM doing arithmetic is a hope, not a guarantee — shot durations are now reconciled deterministically against the requested runtime, because a "15 second" film that runs 23 seconds breaks every downstream cost estimate and render length.

Accomplishments that we're proud of

No AI tool today gives you a complete production package from one prompt. They give you a clip. Worldsmith gives you research with citations, a world bible, a storyboard, continuity-checked footage, a cut, a narration track, and an 8-platform campaign — from a single sentence.

The three things we're most proud of are the ones a generic agent runtime doesn't give you: shot durations reconciled deterministically rather than trusted to a model; continuity QC that can send a shot back rather than filing a report; and credits reserved before a provider is called and refunded when it produces nothing. No free generation, no charging for failures.

What we learned

A 200 OK is not a passing test. The costliest bug we hit returned success, returned well-formed data, and rendered beautifully. It was only detectable by asking the system for something absurd and checking whether the answer changed. We now test integrations by varying the input and asserting the output varies.

Serverless changes what "background work" means. A fire-and-forget promise is not a job. Anything that outlives a response has to be resumable, and every stage needs to persist enough that re-entering is cheap.

Provider seams paid for themselves. Being able to run the entire pipeline in mock mode meant we could develop the orchestration, the QC gate and the cost ledger without spending anything — and it made every real failure obvious, because the only thing that had changed was the provider.

What's next for Worldsmith

We're taking it live. Billing and subscriptions are already integrated, the credit ledger and plan tiers are built, and the product is feature-complete. What remains is a domain and a support address — that's genuinely all.

After that: a feedback loop that closes the circle. Performance signals from published posts feed back into the next discovery cycle, so the research stage learns from what actually worked rather than starting cold every time. Longer runtimes, more shot types, and a shared Cast library so a character built once can carry a series.

Built With

Share this project:

Updates

Submission history