Inspiration
While working on a project that used Gemini Live for a service I support, I started wondering what this technology could become in a more imaginative setting. Instead of using it only for utility, I wanted to apply the same real-time conversational foundation to something magical for children.
Most children’s content is passive. They watch it, but they do not shape it. I wanted to create something different. I wanted to create a story world that listens, responds, and changes with a child’s imagination in the moment.
I was especially interested in the idea that stories can do more than entertain. They can help children build confidence, curiosity, and language skills. A child might bring a favorite toy into the story and let that character do something brave first, like entering a dark cave, meeting new friends, or exploring something unfamiliar. In that way, the experience becomes more than a story. It becomes a playful space where children can practice imagination, hear language in context, and explore courage in a way that feels safe.
The goal was not just to make another AI story app. It was to create an experience where a child can speak an adventure out loud, see it come to life, and feel like they are creating the story instead of just consuming it.
What it does
Voxitale is an interactive AI storybook that turns a child's imagination into a personalized animated movie in real-time.
- A child begins by speaking their adventure out loud. Amelia, the AI narrator, listens and adapts the story instantly, allowing the child to guide what happens next simply by talking.
- As the story unfolds, the system generates illustrated scenes using Gemini Flash 2.5 so the child can watch their ideas come to life visually. If they decide to change the direction of the story, Amelia adjusts the narrative and new scenes are created to match.
- The physical environment becomes part of the experience. Through integration with Home Assistant, lights in the room shift color and pulse to match the mood of the story. A dark cave may dim the room while a magical discovery might fill the room with warm glowing light.
- Children can also bring their own characters into the story. Using a camera, they can capture an image of a toy, drawing, or object and the system will incorporate it into the adventure.
- When the story ends, all generated scenes are assembled into a short animated movie complete with narration, music, and sound effects that the child can watch again and keep as a memory of the adventure they created.
- Behind the scenes, a meta learner runs after each session using evolutionary algorithms to experiment with improving prompts based on human and AI feedback, helping the system gradually improve the quality and coherence of future stories.
How we built it
- Voxitale is built as a real time multi-agent storytelling system powered by Gemini Live and Google Agent Development Kit (ADK).
- At the core of the system is a live storyteller agent built with Google ADK. This agent manages the conversation with the child and coordinates supporting agents responsible for planning scenes, generating illustrations, and reviewing the final story output.
- The voice interface is powered by Gemini 2.5 Flash bidirectional streaming, which enables low latency conversation between the child and the narrator. This allows the child to interrupt, redirect, or expand the story naturally while the system responds in real time.
- The backend runs as a FastAPI service deployed on Google Cloud Run. The storyteller agent manages the live voice session and coordinates scene generation during the story. When the story concludes, a separate Cloud Run job assembles the generated scenes into a final movie using FFmpeg.
- The frontend is built with Next.js 15 using React 19 and TypeScript with the App Router architecture. It handles the live storytelling interface, scene rendering, and communication with the backend session.
- Scene illustrations are generated using Gemini Flash 2.5 as the story evolves. Narration is produced using ElevenLabs text-to-speech (TTS), with Google Cloud Gemini TTS available as a fallback.
- To extend the experience into the physical environment, the system integrates with Home Assistant. This allows lights in the room to react to the story and change color or intensity to match the narrative mood.
- Infrastructure, deployment, and environment configuration are managed with Terraform to support a scalable and reproducible cloud environment.
- Much of the project was developed with the help of Google Anti-Gravity and OpenAI Codex. I still care deeply about understanding how systems work, but my relationship with coding has changed. AI made it possible to move faster and build a more polished project under tight time constraints, even if it sometimes took away some of the old feeling of getting lost in the craft and discovering solutions the long way.
Safety and privacy considerations
- Because Voxitale is built for children, safety and privacy are core product requirements from the start. The experience begins with an adult facing setup screen where a parent or caregiver can choose the story mood, age range, and pacing. Before the story starts, a dedicated microphone check helps create a smoother and more controlled experience.
- I also added child focused guardrails to keep the storytelling G rated, age appropriate, and emotionally safe. The system is explicitly guided away from horror, gore, jump scares, and other frightening themes. On the visual side, the image pipeline avoids visible text and brand logos. If a child references a famous character, Voxitale treats that as toy inspired creative guidance rather than attempting to reproduce protected branded details directly.
- Privacy and family control are equally important. Toy photos are optional and child led, and transcribed speech is scrubbed for personal information before it is logged. I also want to be transparent that this is still an area I am actively improving. Some session data is currently stored to support story continuity, image generation, and final movie assembly, and uploaded toy images are saved as session artifacts rather than treated as purely temporary inputs.
- At this stage, family control mainly comes through adult setup choices, optional sharing flows, and session cleanup or purge paths. In a more mature version of the product, I would add clearer parent managed save and delete controls, stronger retention policies, and better transparency about what is stored, for how long, and for what purpose.
Challenges we ran into
Building a real time storytelling system for children introduced several challenges across networking, media generation, audio, and consistency.
- Google ADK Bidi provided the core live conversational channel, but the experience around it was highly custom. I built my own websocket routing and session layer to coordinate scene generation, image and theater updates, interruption handling, and story state across the UI. That meant reconnects had to do more than restore the audio stream. They also had to rehydrate story context, recover in progress generation or movie assembly state, and replay any still relevant queued updates so the experience felt continuous rather than reset.
- Interruptibility was another major challenge. Children naturally change their minds in the middle of a story, so the system needed to support scene changes while an image was already being generated without breaking the conversation or causing conflicting outputs.
- I also had to deal with duplicated messages coming from Gemini Live, which could create confusing or repeated narration if not filtered properly.
- Image generation speed was a constant challenge. Latency from Gemini Flash 2.5 and later image models could slow the pace of the story, and large generated images sometimes created additional load time issues on the frontend.
- Maintaining visual consistency across a story was difficult as well. Characters, environments, and scene details needed to remain recognizable from one moment to the next even though each image was being generated independently.
- ADK Bidi handled the core live voice channel, but achieving smooth audio quality and natural turn taking still required significant application level logic. I built custom microphone gating, greeting protection, interruption control, playback buffering, and audio recovery around the live session to handle room noise, premature speech pickup, silent turns, and browser audio glitches. I also added watchdog logic to detect and reset background audio whenever it stalled or failed to resume cleanly, so the experience stayed fluid and dependable for children.
- Another challenge was handling branded or copyrighted characters. When children referenced familiar characters, the system needed to fall back to something generic while still preserving the spirit of the story.
- At a systems level, repeated image generation could also create resource strain and model fatigue, especially in longer sessions.
- Finally, synchronizing text-to-speech with highlighted words in the story interface proved difficult, especially when working across different TTS systems and trying to maintain a smooth read along experience.
Accomplishments that we're proud of
- Built interruption handling that allows a child to change the story mid generation without breaking conversation flow or story state.
- Solved numerous real time edge cases to make the experience feel fluid, resilient, and natural for children.
- Improved character and scene continuity by introducing continuity anchors and thumbnail based reference images.
- Integrated Home Assistant lighting so the physical room reacts to the story and becomes part of the experience.
- Preserved immersion across reconnects by persisting session state, buffering outbound events, and restoring still relevant updates instead of resetting the session.
What we learned
- Building a stable real-time architecture around Gemini Live required more custom recovery and self-healing logic than we originally expected.
- In children’s storytelling, latency matters as much as quality because even strong outputs lose value if they arrive too slowly to hold attention.
- Building for children is different from building for adults because delays, awkward turn taking, and inconsistent feedback are noticed immediately.
- Provisioned throughput and token heavy media generation become important scaling concerns in any production version of this experience.
- This type of experience would likely work better as a mobile app, where device control, audio coordination, and connected experiences are easier to manage than in the browser.
What's next for Voxitale
- I want to continue refining Voxitale through more testing so the experience feels smoother, more stable, and more polished for children and families.
- I also want to go deeper on research around character and scene consistency across long form stories, especially as large language models and world models continue to improve. That remains one of the biggest challenges in making the experience feel truly magical and coherent from beginning to end.
- Another priority is adding stronger logging, evaluation, and benchmark story testing so I can better measure improvements in continuity, responsiveness, and temporal coherence across scenes.
- I also want to strengthen family safety, privacy, and accessibility by adding clearer parent managed settings, better retention and deletion controls, stronger transparency around what is stored and why, improved captions and read along support, adjustable pacing and sensory intensity, and alternative interaction paths for children who may not want voice to be the only way to participate.
- On the creative side, I would love to explore Veo as a way to turn these story sessions into richer animated films. In many ways, children are already creating their own storyboard through conversation, and over time that could become the foundation for films that even adults would never think to imagine.
- I also want to make the world feel more immersive through spatial audio and richer sound design so scenes feel more alive around the child, not just on the screen.
- Looking further ahead, I am especially interested in exploring world models such as Genie 3 so these story spaces can become more interactive. My long term vision is not just for children to watch their story, but to step into it.
- Voxitale has been a solo build so far, and one of the clearest next steps is bringing in the right collaborators to help mature it into a stronger product across child experience design, safety, accessibility, and overall production quality.
Built With
- adk
- elevenlabs
- fastapi
- ffmpeg
- firestore
- gemni
- google-cloud
- google-cloud-run
- google-vertex-ai
- home-assistant
- next.js
- python
- react
- tts
- typescript
- websockets

Log in or sign up for Devpost to join the conversation.