-
-
One sentence in, a production out. An autonomous media studio built on Gemini, Veo and Parallel Search.
-
The Studio. Type one sentence, pick a look and a runtime, press Build my world. Everything after that is autonomous.
-
Research runs first, here against a demonstration corpus. The World Bible is built from what it found.
-
A different idea is a different search: the objective comes from your prompt, and each run carries its own search_id.
-
The World Bible: premise, visual style and a character sheet. Every downstream stage obeys these rules.
-
A second production from a different sentence. The research changed, so the world changed with it.
-
The shot list. Each shot carries its own continuity contract, with durations reconciled to the runtime you asked for.
-
Character references and environment plates, each costed per asset as it is produced.
-
QC is a gate, not a report. A frame is rejected for showing clock faces where the World Bible specified glowing amber lenses.
-
Expensive steps sit behind human approval gates. Nothing renders until the frames it depends on have passed.
-
Continuity QC on finished video, not just stills — every clip checked against the World Bible before assembly.
-
The finished 15-second film, assembled with FFmpeg and narrated with Gemini TTS.
-
On-model campaign creatives, generated from the same World Bible that produced the film.
-
One film becomes a campaign: titles, descriptions, hashtags and creatives for eight destinations.
-
Every provider resolved live: Reasoning, Image, Video, QC, Narration on Vertex — Research on Parallel.
Inspiration
Every AI video tool gives you a clip. None of them give you a production.
A real studio doesn't start with a prompt. It starts with research — what does the audience actually want right now? Then it builds a world, keeps every character on-model across every shot, quality-checks the footage, assembles it, narrates it, and ships a campaign.
That whole loop is the bottleneck. We wanted to find out whether a network of agents could hold creative context from the first market signal all the way to the final post — not generate a clip, but run a studio.
What it does
You type one sentence. A network of agents then:
- Discovers — Parallel Search scans live web signals for a real content opportunity
- Thinks — Gemini turns those signals into a content angle and creative strategy
- Creates — builds a persistent World Bible: characters, locations, visual rules
- Produces — frames → Veo video → VLM continuity QC → FFmpeg assembly
- Distributes — an 8-platform campaign with on-model creatives and copy
Expensive steps sit behind human approval gates, and every asset is costed in a transparent per-asset ledger before a cent is spent.
Beyond the pipeline, 11 standalone tools work on their own — Text→Image, Text→Video, Image→Video, Voiceover+Images→Video, Text→Speech, Image→Prompt, Upscale, Social Post, YouTube Kit, Creative Text Editor, and Cast (build a character once, reuse it in any scene).
How we built it
Five planning agents hand each other schema-validated objects, not prose, so a later stage can rely on the shape of what it receives:
| Agent | Produces |
|---|---|
research-agent |
Evidence gathered through Parallel Search, carrying its sources |
opportunity-agent |
The angle worth making, argued from that evidence |
world-builder-agent |
The World Bible — characters, locations, props, visual language |
storyboard-agent |
Shots, with durations reconciled to the requested runtime |
production-planner-agent |
Per-asset model and cost plan |
Every model in the system is a Google model — Gemini on Vertex AI for reasoning,
gemini-2.5-flash-image for stills, Veo 3.1 for video, Gemini VLM for continuity
QC, Gemini TTS for narration. Firebase handles auth, Firestore and Storage.
There are no third-party AI providers anywhere in the stack.
QC is a gate, not a report. Every generated shot goes back to a vision model holding the World Bible and is asked whether the character, location and look actually match. A pass moves forward; a fail regenerates that shot against the same reference, bounded so a stubborn shot can't burn the budget.
Every provider sits behind an interface chosen by one environment variable, so the entire pipeline runs end to end with zero API spend. That's how it was developed.
Challenges we ran into
The Parallel integration looked like it worked, and didn't. We were sending
the planned queries as { query: "..." }. The API's field is search_queries —
an array. Parallel ignores unknown fields rather than rejecting them, so every
request returned HTTP 200 with plausible, well-formed results. They were just
ranked against our hardcoded objective alone, identically for every run. We
found it by searching something deliberately absurd: asking for "competitive axe
throwing league Estonia" returned articles about social media trends. The same
key sending search_queries returned axe-throwing venues in Tallinn. A second
bug hid under it — result bodies arrive in excerpts, an array, and we were
reading snippet, so every piece of evidence carried an empty body and the
synthesis step was reasoning over bare titles while appearing to cite sources
correctly. Both were silent. Neither would ever have thrown.
Fixing the research broke the pipeline. With real excerpts in the prompt, synthesis went from 6 seconds to 36. The six-call planning chain stopped fitting inside a serverless function's execution limit, and the function was terminated rather than throwing — so nothing was ever written as an error, and the UI waited forever on a status that could never change. Every stage already persisted its own output, so we made each one skip itself when its output exists and added a resume path the client triggers when a run stops advancing. Completed work is never re-run, so a resume can't duplicate spend.
One credential resolver served two clouds. When we moved Vertex to a separate Google Cloud project from Firebase, the shared service-account resolver preferred the Vertex credential — and it was also used by Firebase Auth, Firestore and Storage. The project id still came from the Firebase env var, so nothing looked misrouted; sign-in, credits and every asset write would simply have started failing against a project that account had no permissions in. Model plane and data plane are now separate resolvers with opposite precedence.
firebase-admin reaches jose through a CommonJS require. On Node below
22.12 that throws ERR_REQUIRE_ESM — and the app otherwise appears to start
normally while every route touching Auth or Firestore returns a blank 500 with a
minified React error. We pinned the Node floor and added a health endpoint that
actually initialises Firebase Admin rather than just parsing the key.
Smaller ones that each cost a day: signed URLs expire, so a 7-day URL persisted into a database is a 404 with a delay fuse — assets now resolve through a route. Data URLs above a few hundred KB blow the server-action argument limit, which silently broke every image-taking tool for real photographs. The cost ledger is read-modify-write, so a retry mid-run clobbered the run's own accounting until we serialised per-project writes. And an LLM doing arithmetic is a hope, not a guarantee — shot durations are now reconciled deterministically against the requested runtime, because a "15 second" film that runs 23 seconds breaks every downstream cost estimate and render length.
Accomplishments that we're proud of
No AI tool today gives you a complete production package from one prompt. They give you a clip. Worldsmith gives you research with citations, a world bible, a storyboard, continuity-checked footage, a cut, a narration track, and an 8-platform campaign — from a single sentence.
The three things we're most proud of are the ones a generic agent runtime doesn't give you: shot durations reconciled deterministically rather than trusted to a model; continuity QC that can send a shot back rather than filing a report; and credits reserved before a provider is called and refunded when it produces nothing. No free generation, no charging for failures.
What we learned
A 200 OK is not a passing test. The costliest bug we hit returned success, returned well-formed data, and rendered beautifully. It was only detectable by asking the system for something absurd and checking whether the answer changed. We now test integrations by varying the input and asserting the output varies.
Serverless changes what "background work" means. A fire-and-forget promise is not a job. Anything that outlives a response has to be resumable, and every stage needs to persist enough that re-entering is cheap.
Provider seams paid for themselves. Being able to run the entire pipeline in mock mode meant we could develop the orchestration, the QC gate and the cost ledger without spending anything — and it made every real failure obvious, because the only thing that had changed was the provider.
What's next for Worldsmith
We're taking it live. Billing and subscriptions are already integrated, the credit ledger and plan tiers are built, and the product is feature-complete. What remains is a domain and a support address — that's genuinely all.
After that: a feedback loop that closes the circle. Performance signals from published posts feed back into the next discovery cycle, so the research stage learns from what actually worked rather than starting cold every time. Longer runtimes, more shot types, and a shared Cast library so a character built once can carry a series.
Built With
- dodo-payments
- ffmpeg
- firebase
- firebase-auth
- firebase-storage
- firestore
- gemini
- gemini-tts
- google-cloud
- google-genai-sdk
- next.js
- node.js
- parallel-search
- react
- tailwindcss
- typescript
- veo
- vercel
- vertex-ai
- zod
Log in or sign up for Devpost to join the conversation.