Inspiration
Every AI video tool has the same failure mode. You generate a shot, a character's face has drifted or a phantom duplicate has appeared at the frame edge, and the only button you have is "generate again". You press it. Sometimes it works. Nobody records why it failed, so nobody learns anything, and retries quietly become the largest line item in the budget.
Narrative animation makes this worse than it is for single clips, because consistency is the whole product. A character has to be the same character in shot 1 and shot 40. A world has to stay the same world.
So we built the thing that was missing: a studio with a production memory. Not a better prompt — a system that knows what went wrong last time and what actually fixed it.
What it does
Chalchitra takes a story premise and runs it through a crew of Gemini agents to a delivered cut:
- Kathakaar writes the script — 3-5 beats, 1-3 scenes, sized for a two-minute short
- Roopkaar designs the cast as checkable parameters (face shape, body type, hair/fur shape and colour, outfit, signature accessories, and a scale note), never as adjectives — because "a small brown owl with a red scarf" drifts, while a parameter list can be restated identically in every prompt and verified literally from a still
- Kalakaar locks one style key for the whole film, then writes each scene's set reference in that style, plus a cast lineup that fixes relative body sizes
- Sutradhaar breaks the script into 4-8 second tiered shots. The duration cap is architectural: long takes are where identity drift and physics breakdown live, so the studio only ever generates short shots and stitches them
- Drishti inspects generated footage against the locked references and scores six dimensions — cast, identity, scale, accessories, background, style. The overall score is the worst dimension, never an average, because a film is only as consistent as its worst visible violation
- Vichaarak is the one that makes this a studio rather than a retry loop
Cast, identity, and scale are treated as blocking gates, not warnings. A phantom duplicate character or a broken size relationship stops delivery.
When a shot fails, Vichaarak connects to the studio's ClickHouse cluster through the official ClickHouse MCP server and investigates with real SQL: how has this shot behaved across attempts, do similar shots fail the same way, is this specific or systematic, and which remediation action has actually improved this class of failure before. It returns a structured decision — failure class, action, confidence — with the SQL it actually executed attached, so a director can re-run every query themselves.
The evidence trail is captured from observed tool calls, not from the model's description of what it did. The agent cannot inflate its own homework.
In our seeded production history the patterns are unambiguous: cast failures occur on 58% of two-character shots and 0% of solo shots, and regenerating with a fresh validated keyframe improved 6 of 6 blocked shots (average +0.49 overall) while plain regeneration improved 2 of 4 (average +0.05). That is the difference between spending money and learning something.
How we built it
Agent runtime. Google Agent Development Kit. Each agent is an LlmAgent with a
Pydantic output_schema, so a model that drifts produces a validation error
instead of a broken film. Structured-output agents are kept tool-free and the
MCP-holding analyst is paired with a tool-free decision formatter, which keeps the
design portable across ADK versions.
Models. Gemini 3.1 Pro for script writing and evidence analysis, Gemini 3.8
Flash for high-volume structured planning and vision QC, Nano Banana Pro
(gemini-3-pro-image) for reference sheets and keyframe composition because it
accepts many reference inputs at once, Veo for video.
Partner integration. The official ClickHouse MCP server (mcp-clickhouse),
launched as a stdio subprocess and attached to the analyst as an ADK McpToolset,
with Streamable HTTP supported for deployed use. The server stays in its default
read-only mode, so "the analyst can only read" is enforced by the server rather
than by prompt instructions.
Production memory. ClickHouse, append-only across five tables — generation
attempts, machine continuity verdicts, human director labels, agent decisions, and
pipeline events — plus five analytical views the agent leans on: failure_clusters,
remediation_outcomes, judge_alignment, cost_rollup, latest_shot_scores.
Corrections arrive as new rows, which is what makes "did our fix actually work?"
answerable at all.
Rest of the stack. FastAPI, Firestore for mutable project state, Cloud Storage for media, ffmpeg for frame extraction and assembly, Cloud Run for hosting.
Challenges we ran into
A genuine dependency deadlock. google-adk requires mcp 1.x; mcp-clickhouse
requires mcp 2.x. They cannot share a virtualenv. We run the MCP server in an
isolated uvx environment with a pinned version — which is the launch pattern
ClickHouse documents — so both sides stay on supported versions and the
integration is still a real runtime MCP connection.
Making an agent's evidence trustworthy. An agent that says "identity fails on two-character shots" is worthless if you cannot check it. We capture tool calls from the run's event stream and store the executed SQL alongside the decision, and we clamp confidence when the agent ran no queries at all.
Keeping policy out of the model's hands. The retry budget, the blocking thresholds, and the exact-cast rule are enforced in code after the model replies. An agent that would prefer to try once more on an exhausted shot gets overruled and the shot goes to a human.
Scoring honestly. Averaging the six QC dimensions hides exactly the failures that matter, so the overall score is the minimum. And cast is not a threshold at all — it is an exact-match test, because there is no such thing as 70% of the right cast.
Accomplishments that we're proud of
- The remediation loop is closed: a failure is classified, a fix is chosen from measured history, the fix is applied, and the outcome is recorded so the next decision is better informed
- Every agent decision is auditable down to the SQL that justified it
- The consistency machinery is mechanical rather than prompt-based — parameter sets, a locked style key, a cast lineup for scale, and a validated exact-cast keyframe used as the video's literal first frame
- Blocking gates that actually block. Export is refused, not warned about
- Human labels live next to machine verdicts, so you can tell when the threshold is wrong rather than the film
What we learned
Consistency is an accounting problem before it is a generation problem. Once you record every attempt and every verdict, the failures stop looking random and start clustering — by character count, by provider, by scene, by which reference was never locked. You cannot fix what you have not grouped.
ClickHouse turned out to be the right shape for this, and not for performance reasons. Its append-only grain matches how production actually works: you never edit a past take, you shoot another one. That constraint is what makes the history trustworthy.
We also learned how much of "agentic" work is knowing what not to let the agent decide. The interesting judgement is diagnostic. The rules — retry budgets, hard gates, cast exactness — belong in code.
What's next for Chalchitra
- Continuity across chapters, so a series keeps one cast and one world
- Learned routing: pick the model per shot from measured success rates for that shot's character count and tier, instead of a static tier policy
- Turning
judge_alignmentinto automatic threshold calibration from director labels - Native audio and score synced to cut points
- Cost forecasting before a render starts, from historical retry rates for comparable shots
Built With
- clickhouse
- cloud-run
- cloud-storage
- docker
- fastapi
- ffmpeg
- firestore
- gemini
- gemini-3.1-pro
- gemini-3.8-flash
- google-adk
- google-cloud
- mcp
- model-context-protocol
- nano-banana-pro
- pydantic
- python
- veo
- vertex-ai
Log in or sign up for Devpost to join the conversation.