Inspiration
Reading takes mental effort. Many people struggle to visualize complex written worlds. Meanwhile, AI video is booming, but it hallucinates constantly. Most generated clips contradict the source material. Studios sit on thousands of unmade books because pre-visualization is too expensive. We realized the missing link isn’t a bigger video model. It is structured data. We built ShotBook to ground cinematic generation in actual literary canon.
What it does
ShotBook turns any ebook into an interactive, visual platform. A user highlights a passage. The system instantly generates a cinematic scene true to the story. It does not guess. It tracks character descriptions, evolving personalities, and shifting settings across the entire text. It gives everyday readers an imagination boost. It gives film studios a zero-cost pre-visualization engine.
How we built it
We built a two-stage pipeline:
- The Ingestion Pass: We break a book into paragraphs and feed it to Qwen 2.5 32B running across 8 H100s. The model extracts baseline character profiles, location lore, and ambient details into a PostgreSQL database. A second pass captures temporal "deltas" (e.g., Character A is now wounded) and maps them to specific paragraph intervals.
- The Generation Engine: When a user highlights text, our backend queries the database. It resolves the "baseline" lore against active "deltas" to lock in the exact scene reality. It compiles this into hyper-specific shot prompts. Finally, Wan 2.2 generates the 5-second clips and stitches the sequence together.
Challenges we ran into
State tracking is notoriously fragile. Reconciling a character’s "static baseline" with a "temporary delta" required writing strict programmatic conflict-resolution logic. Our biggest roadblock was the audio layer. We built a pipeline using FishSpeech for dialogue and AudioGen for sound effects. However, dynamically mixing them and frame-syncing them to Wan 2.2's output proved too mathematically unstable to ship in a weekend.
Accomplishments that we're proud of
We solved the "continuity problem" in generative video. Instead of lazy prompting, we built a deterministic state-machine for literature. We successfully parallelized a 32B LLM to parse an entire book's subtext without losing context. Most importantly, we proved our thesis: better data context beats a heavier video model.
What we learned
Single-pass LLM extraction misses nuance; a "baseline-then-delta" architecture is mandatory for long fiction. We also learned that video models desperately want to be put in a box—the tighter the spatial constraints we fed Wan 2.2, the higher the cinematic quality. Finally, multi-modal audio syncing cannot be an afterthought; it requires its own dedicated temporal pipeline.
What's next for ShotBook
- Frame-to-Frame Continuity: Passing the final frame of Clip A into Wan 2.2 as the reference image for Clip B.
- An Evaluator Agent: Slicing generated video into keyframes, running an image-captioner over them, and mathematically scoring the visual output against our Postgres "ground truth" database.
Built With
- asyncio
- asyncpg
- audiocraft-audiogen
- beautiful-soup
- bore.pub
- claude-api-/-anthropic-sdk)-data/context-pipeline-(postgresql
- claude-api-/-anthropic-sdk-postgresql
- cuda
- eslint)-corpus-ingestion-(pdfminer.six
- eslint-pdfminer.six
- ffmpeg
- fish-speech
- git
- httpx
- hugging-face-diffusers
- hugging-face-hub
- imageio
- llama-3-70b-instruct
- make
- numpy
- ormsgpack)-frontend-(react
- ormsgpack-react
- pillow
- pydantic)-backend-api-(fastapi
- pydantic-fastapi
- pytorch
- react-router
- requests)-infra-&-tooling-(8x-h100
- requests-8x-h100
- sqlalchemy
- supabase
- tailwind-css
- tmux
- triton
- typescript
- uvicorn)-audio-pipeline-(llama-3.1-8b-instruct
- uvicorn-llama-3.1-8b-instruct
- vite
- vllm
- wan-2.2-t2v
- websockets

Log in or sign up for Devpost to join the conversation.