💡 Inspiration
A screenwriter has a 45-minute script on her desk. She has never written a prompt in her life.
That isn't a gap she's going to close. She spent those years learning to write, not learning which adjectives a diffusion model happens to respond to. And every tool that could film her story assumes she already speaks its language: which quality tier, which aspect ratio, what a "reference image" is for, how much three minutes of footage is going to cost her before she commits.
Nobody hands her any of that. So the film doesn't get made.
There are far more good scripts in the world than there are people who know how to write a prompt. That gap is the whole problem, and it isn't really a technical one. It's a translation problem.
Most people can't really talk to AI. We built the thing that stands in between.
We made that testable rather than leaving it as a slogan. If a text box shows up anywhere in the product besides file upload, we've broken our own rule. When something is missing, the system makes candidates and lets her pick. It never hands her an empty field and asks her to describe what she wants.
In this business a film gets made when someone green-lights it, and that decision has always belonged to whoever holds the money. We wanted it to belong to the person who wrote the thing.
🎬 What it does
She drags the file onto the page.
Ninety seconds later she's looking at 37 scenes, 6 characters, and a sentence no tool has ever said to her before: this chapter is 12 shots, it'll cost $4.50, and about half of them will come out right the first time.
She picks a visual style. The buttons themselves tell her photorealistic is ten points harder than the anime option next to it, so she picks anime, and the estimate updates while she watches. She casts her characters from generated candidates, three at a time, two cents each. Nothing has been filmed yet. She has spent forty cents.
Then she presses one button and watches twelve shots render, one at a time, each one waiting for her to say keep or again. When it's done there's one file, subtitles burned in, in whichever of three languages she chose.
She never typed anything except the script.
That middle part, the part she doesn't know how to do, is the entire product. A few things about how it behaves are worth explaining.
Casting is a hard gate, and we can show why. No shot generates until every character in it has an approved reference image. Same shots, the only difference being whether that reference existed:
| First-pass success | Expected retry cost | |
|---|---|---|
| Before casting | 29% | $1.86 |
| After casting | 49% | $1.17 |
Twenty points, for two cents. That table is the argument for the cost ladder the whole product sits on: a candidate image is $0.02, a shot is $1, a finished segment is $20. Wherever there's a decision to hesitate over, we try to move it to the cheap end.
She can be in her own film. Drop a photo of herself onto the character card and the designer treats it as the face, not the card: it comes back with three versions of her in the film's style, and every later shot she appears in is her. When chapter seven needs a different outfit, she opens the character, picks chapter seven, and asks for a new wardrobe — the face is locked, only the clothes move, and each chapter keeps its own look. Before any upload is accepted she ticks one box: this image is hers to use, and the liability is hers.
Style changes which physics matter, not just how things look. Photoreal weights hand anatomy at 1.3×; 2D animation weights it at 0.5×, because over-constraining hands in a cel-shaded frame just fights the art style. That's the uncanny valley turned into a setting.
It looks at every clip before she does. The moment a shot finishes, a Gemini model is handed the actual footage and asked for observable facts: how many people are in frame against how many the shot called for, whether text has crept onto the screen. The findings sit on the shot card when she decides. And when she presses again, the rejected take is written to the ledger first, findings attached, and the retry prompt carries them — the previous take showed two people; this shot has one. A reshoot isn't the same dice rolled twice. The loop closes between one take and the next, not at the end of the project.
Dark dialogue is normal in drama, and safety filters don't know that. When a line trips the content filter we don't ask her to rewrite her screenplay. We generate the shot quietly and put the line in the subtitle layer instead. Not one word of her script changes.
And she can take the prompts and leave. After casting, every prompt is exportable, as plain text, as Markdown with the reasoning attached, or as JSON. We figured a tool that's willing to hand your work back is one you can trust with it.
🗄️ Why ClickHouse is the part that makes it honest
The most useful sentence this product produces is "about half of these will come out right on the first try."
Without history, that's just a confident-sounding number.
So every generation writes one row: what kind of shot it was, the parameters, the quality tier, which attempt this was, whether the take got used, what it cost, how long it took, which model. When someone plans a similar shot later, we query that history and compute a real pass rate instead of a prior, and we label which of the two you're looking at, together with the sample size.
It runs through the official mcp-clickhouse MCP server. Two materialised views do the actual work:
| View | The question it answers |
|---|---|
constraint_stats |
For each physical constraint, what pass rate do we actually observe? |
calibration_stats |
Are our predictions systematically optimistic or pessimistic? |
The second one exists because a success-rate predictor nobody checks is just a number that sounds sure of itself. During testing we seeded a deliberate ×0.92 bias into the ledger; the report caught it and said 6.3 points optimistic.
Three things we say out loud rather than tuck away:
- Cold start is labelled. On day one the ledger is empty, so every estimate says "prior estimate, no observations yet," plus how many samples it takes before it switches over.
- The observed constraint lift isn't causal, and the report says so. Constraints fire based on how hard a shot is, so harder shots get more of them. That's selection bias, and we'd rather name it than quote a number we can't defend.
- Backfilled rows are flagged and left out of calibration. Their predictions were recomputed with priors that had already been tuned against those same outcomes. Counting them would be marking your own exam with the answer key open.
All of it is queryable from the UI. A statistic can be made up; timestamped rows carrying model ids, real costs and real latencies can't. For a partner track that's the difference between saying we use ClickHouse and being able to show it.
Once a style has enough observations, the system distils the ledger into a short note in its own words and carries it into every later storyboard. The first one it wrote for photoreal, from 39 first takes, told us two things we had assumed wrong: frame-chaining makes no real difference to pass rate (78% vs 75%), and our own estimates run conservative — shots it predicted at 20–50% actually passed 67–100%. We didn't write those sentences. The ledger did.
If ClickHouse is unreachable, planning falls back to priors, keeps working, and the UI says so. The dependency is real, but it isn't a single point of failure for her.
🔧 How we built it
One Cloud Run service. Frontend and backend ship together, so there's one URL and one deploy. Gemini on Vertex AI throughout: the heavier model reads the screenplay and plans shots; Flash writes prompts, translates subtitles and audits finished footage; Flash-Lite handles the high-volume batch work. Gemini Image for casting, Gemini Omni Flash for video, Gemini Flash TTS for narration. Firestore for state, Cloud Storage for media, ClickHouse for the learning loop.
One decision worth naming: the gates are real state transitions, not requests inside a prompt. Every action that costs money changes state in the schema. We never ask the model to please confirm before spending. Probabilistic output can't guarantee a hard requirement, so whatever the architecture can enforce, the architecture enforces.
💸 Challenges
Our budget was $100. Ten seconds of video costs a dollar.
So we never got to do the thing a funded team does when a shot comes out wrong, which is run it again. Every press had to be worth it. The full film prices out around $238, and we still haven't shot all of it.
But that's the reason the product exists at all.
Nobody with a budget builds success-rate prediction. If it's wrong, run it again. Nobody with a budget builds a cost ladder, or writes every failed take into a database so the next estimate is a little better. There's no reason to.
And being broke is exactly where our screenwriter is standing. She doesn't get to "just generate it again" either.
We were broke, so we built the thing broke people need.
It became a real dimension of the product rather than a complaint. Every quality tier is priced in config, and the system quotes her selection before she commits. Bigger budget, higher tier and resolution. Smaller budget, a cheaper tier, and she still finishes the film. Same script, she can shoot it cheap to see whether the story lands, then re-shoot a handful of moments at the top tier.
Whether the tool is expensive was never the point. The point is that she knows the number before she presses the button. The usual arrangement is generate first, see the bill later. Ours is quote, decide, then spend.
📚 What we learned
A 200 doesn't mean it worked. Subtitle burn-in returned 200, produced a 30MB file, and had no visible text, because the container had no fonts and ffmpeg exits 0 regardless. Our generation timeout was set to 300 seconds, but requests measures the gap between reads rather than total elapsed time, so a connection that stays open and never answers never trips it; the request ran until the platform killed it at 900. Being killed is worse than timing out, because the exception handler never runs, so the clip state and the failure record are both gone. Our rule now: degrading is fine, degrading silently is not.
A check that covers half the problem is worse than no check. We wrote a health endpoint that reported "12 fonts, 4 CJK." It looked perfect, and we cited it as proof. But fontconfig's config was missing, so libass still drew nothing. It handed us a green light that looked like someone had verified it.
Two implementations of the same thing will drift. The homepage and the project page each had their own learning-loop component; one read a field under the wrong key and quietly showed 0 for three days. They're one component now.
She describes the symptom, not the cause, and the symptom is always real. Every "it's stuck" came with a diagnosis that turned out to be wrong, and not once did the report itself point at the wrong place.
🚀 What's next
Finishing the film. That needs budget more than it needs time.
Every iteration aims at the same thing: push the displayed success rate up, so it costs her less to get what she was after. Taking that barrier down is really the only job we have.
This is version one. It keeps changing, because we use it ourselves.
Built With
- clickhouse
- cloud-run
- cloud-storage
- fastapi
- ffmpeg
- firestore
- gemini
- google-adk
- google-cloud
- google-genai
- model-context-protocol
- pydantic
- python
- react
- typescript
- vertex-ai
- vite

Log in or sign up for Devpost to join the conversation.