Inspiration
A storyboard can capture a beautiful frame, but it cannot show how that frame should move. Directors still need to communicate character actions, camera movement, atmosphere and pacing before committing time and budget to video generation.
We built FramePilot to bridge that gap. It lets creators step into the director’s chair: upload a storyboard, add a screenplay excerpt and creative intent, and turn that static frame into a grounded motion plan, an instant local preview and, when approved, a cinematic generated scene.
What it does
FramePilot follows a three-stage workflow:
Direct Scene The agent studies the storyboard together with the screenplay and creative intent. It identifies characters, objects, environmental elements, relationships and camera opportunities, then creates an evidence-grounded direction plan.
Motion Preview FramePilot produces a fast local 2.5D preview so the creator can evaluate composition, camera movement and atmosphere without spending a paid video-generation request.
Generated Video If the requested action requires real character, facial, cloth or object movement, FramePilot routes it to a protected generative-video workflow. The creator reviews the exact configuration and explicitly approves the request before Veo is called.
The result is a creative agent that does more than generate video: it helps users reason about motion, preserve the storyboard’s identity and decide when expensive generation is actually necessary.
How we built it
FramePilot was built and deployed using Replit Agent.
The application uses Gemini 2.5 Flash through Vertex AI for multimodal scene understanding. Gemini examines both the uploaded image and the written direction, then returns structured semantic information about:
- Main characters and their intended actions
- Visible objects and props
- Environmental motion
- Camera direction
- Grounding evidence and preservation constraints
A Python validation layer converts this analysis into a bounded motion plan. It rejects unsupported actions, separates characters from objects and environmental elements, limits excessive model output and retains valid fields even when another field needs correction.
The Motion Preview runs locally using a lightweight 2.5D system. This gives creators an immediate, zero-generation-cost way to test camera movement and environmental depth.
For the final video stage, FramePilot uses Veo 3.1 on Vertex AI with a validated storyboard image. Generated outputs are written to a private Google Cloud Storage location and securely served through the application. The job system records provider operations immediately, resumes polling after refresh or restart and restores completed videos without submitting a duplicate request.
Safety and cost control
Generative video can be expensive and unpredictable, so approval is part of the product architecture rather than a warning added afterward.
FramePilot distinguishes between:
- Camera and atmospheric movement that can be previewed locally
- Semantic movement that requires generative video
- Unsupported changes that could distort the original scene
Every Veo request requires an explicit approval step. Authorizations are one-time, scene-bound and session-bound. Duplicate clicks cannot produce duplicate submissions, and refreshing the browser cannot silently generate another video.
This protects both the creator’s visual intent and their cloud budget.
Challenges we faced
One major challenge was making structured multimodal output reliable. Early model responses occasionally contained oversized fields or incomplete JSON. Instead of weakening validation or adding scene-specific rules, we built generic field-level recovery, bounded string and list handling, compact output instructions and a larger structured-output budget.
Another challenge was correctly classifying complex images. A group of people could sometimes be treated as an environmental phrase rather than distinct characters. We improved the semantic boundary so multiple agentive characters remain separate while props and atmospheric elements retain their correct categories.
Veo also introduced lifecycle challenges. A generation operation could continue beyond the application’s original polling deadline even though Google had accepted it successfully. We redesigned the workflow around durable operation IDs, resumable polling and a longer deadline. This allowed FramePilot to recover a completed video rather than incorrectly treating it as permanently failed.
Finally, browser refreshes revealed how important session-safe recovery is. We added durable scene linkage so the correct storyboard, screenplay, intent and completed video return together without mixing outputs from different projects.
What we learned
Building FramePilot taught us that a dependable creative agent needs more than a strong model prompt. It needs clear boundaries between interpretation, validation, preview, generation, approval and recovery.
We also learned that the best way to control generative-AI cost is not simply to reduce usage—it is to give users a useful intermediate result. Motion Preview lets creators make decisions before requesting a paid render, while Veo remains available for moments that genuinely require generative movement.
What’s next
Our next step is to expand FramePilot into a multi-shot workspace with reusable character continuity, shot sequencing and collaborative reviews. We also want to explore richer depth estimation, editable motion paths and comparison views between the original storyboard, local preview and generated result.
FramePilot’s larger vision is simple: make cinematic direction more understandable, controllable and accessible before the expensive part of production begins.


Log in or sign up for Devpost to join the conversation.