Inspiration
A novel is eighty thousand words; a short film is two minutes. Every AI video tool we tried closed that gap by throwing away the story to chase flashy clips. We wanted the opposite: keep the story — its arc, its mood, its characters — and turn a whole book into a real short film. Not a trailer of stock footage, not a slideshow of captioned stills, but genuine moving video anyone could make from a manuscript in the time it takes to grab a coffee, for free and with no login.
What it does
BookToMovie takes an entire novel — a .txt, .docx, or .pdf — and returns a finished two‑minute movie: real generated video clips, a single narrator voice, and an original score, assembled automatically into one MP4. The core idea is aggressive condensation: the whole book is compressed to roughly fifteen story beats told by a narrator over the visuals. That one choice sidesteps the hardest, most expensive parts of filmmaking — per‑character voice casting and lip‑sync — while still producing real video. An advanced panel lets you direct: movie length, background‑music volume, visual style, shot count and length, narrator voice, and swappable models.
How we built it
A single straight‑through pipeline runs the moment you press start. A language model reads the manuscript, condenses it into a title, logline, and beats, then plans a validated shot list with a line of narration per shot. It designs a style frame and character portraits as visual anchors, films each shot as real image‑to‑video, records the narration with text‑to‑speech, and scores one instrumental music bed. FFmpeg does the final cut — crossfades, music ducked beneath the narration, loudness normalized, exported as a self‑contained MP4. Three kinds of AI do three jobs: an LLM writes and plans, a generative media engine produces everything you see and the music you hear, and a TTS voice narrates. Every provider sits behind a swappable adapter, and a mock mode runs the whole pipeline offline for free testing.
Powered by Google's Nano Banana image model, BookToMovie turns each story beat into a richly detailed, on‑style frame — consistent characters, coherent scenes, and cinematic lighting drawn straight from the text. Gemini does the writing and planning, reading the entire novel and condensing it into a tight sequence of story segments, each with its own shot direction and a single line of narration. And Gemini carries that story through to the soundtrack, shaping the narration script that becomes the film's spoken audio — so one model reasons about the book while Nano Banana gives it a face.
Challenges we ran into
Generative pipelines fail in a hundred small ways — a model times out, a clip returns empty, a service hiccups — so our biggest challenge was making the whole thing always finish. We built graceful degradation into every stage: a failed shot holds a still frame and keeps going; a failed score plays silent; a broken plan stops before any paid video work so money is never wasted. Keeping the visuals coherent was just as hard. Early cuts drifted between shots and, worse, seeded wide "establishing" shots from a character's face — so a spaceship over a planet tried to morph out of a portrait. Aligning narration, music, and clips of varying length into one clean timeline took real care in FFmpeg.
Accomplishments that we're proud of
It works end to end, for real: upload a book, get back an actual film with accurate narration and a looped, ducked score. The job never hangs — you always get a movie or an honest error. We shipped a genuine creative surface, not just a demo: sliders for movie length and music volume, free‑text visual style woven into every shot, and per‑job model swaps. We fixed the visual drift by threading the film's style and each character's description into every clip prompt and by seeding establishing shots from a neutral scene frame instead of a face. And the provider abstraction means adding a new vendor is one adapter and one line.
What we learned
The winning move was subtraction. Leaning on a narrator instead of solving lip‑sync is what made a two‑minute film from a whole novel actually feasible. We learned that in generative systems, reliability is a feature — graceful degradation and failing before you spend money matter as much as the output. We learned how much the seed image steers an image‑to‑video clip, and that carrying global style and character context into every prompt is what keeps a film feeling like one film. And we learned to keep secrets on the server with a single ingress point, so keys never leak to the client.
What's next for BookToMovie
Sharper visuals, with shots that track the story even more closely and stronger character consistency across scenes. Longer cuts and pacing presets for different lengths and genres. More voices and multi‑language narration. Optional character casting for users who do want distinct voices. Faster runs through parallel shot generation, a shareable gallery of finished films, and cleaner housekeeping so nothing lingers on disk. Ultimately: from any book to a watchable film, in a click.
Log in or sign up for Devpost to join the conversation.