Inspiration
The name RAWI (راوي) means storyteller in Arabic, an ancient role that turned information into art. We realized that while education technology has advanced, the delivery remains dry—mostly PDFs and static slides. We wanted to bring the "Storyteller" back to the modern classroom using AI, making complex subjects as engaging as a beautifully produced documentary or explainer video.
What it does
RAWI is an AI-powered educational engine that transforms text-based curricula into immersive multimedia experiences. A teacher can input a dry lesson—like the history of Ancient Egypt or the laws of physics—and RAWI automatically generates:
- A structured, age-appropriate educational script with distinct key points.
- High-fidelity infographics, diagrams, and motion graphics that visualize the core concepts.
- Emotive AI voiceovers that guide the student through the journey, making learning multisensory and accessible.
- An interactive AI Chat Assistant alongside the video, allowing students to ask contextual questions about the generated video in real time.
How we built it
We architected RAWI as a multi-agent system using the Google Agent Development Kit (ADK) to orchestrate a seamless flow between curriculum breakdown and media generation.
Brain & Orchestration: We used the ADK to build a specialized "Director Agent" that manages the workflow. It calls Gemini 3.1 Flash to decompose the topic into narrative segments, extract concise "key points" for subtitles, and design detailed visual prompts.
Visual Storytelling: We utilized Imagen 3.0 Fast to produce clean, professional educational infographics and diagrams. We then used Veo 3 to generate cinematic motion graphics that bring those visuals to life. To ensure visual continuity between sequential video segments, we engineered the prompts so that each new Veo prompt explicitly references the context of the previous segment.
Expressive Voices: We integrated the Gemini TTS model to deliver the narrative, creating immersive voiceovers that perfectly pace the educational material.
Media Processing: We implemented an FFmpeg-based pipeline (VideoMerger) with a robust multi-tier fallback mechanism. This pipeline smoothly crossfades video segments, mixes in the voiceover audio, and overlays the generated "key points" as styled, readable subtitles.
Frontend & Backend: The entire backend is built with FastAPI and uses Server-Sent Events (SSE) to stream granular progress (planning → storyboarding → generation → merging) to a dynamic React frontend. It's fully containerized and deployed on Google Cloud Run using a multi-stage Docker build.
Challenges we ran into
Narrative Flow & Continuity: Ensuring the AI didn't just generate random, disjointed videos. We solved this by passing the narrative context of the previous segment into the prompt of the next segment, so Veo 3 maintains topical continuity.
Subtitle Overload: Initially, the system burned the entire verbose narration into the video, which overwhelmed the viewer. We refined the Gemini 3.1 Flash prompt to extract short, scannable "key points" and only burn those into the video as professional semi-transparent subtitles.
Media Synchronization & Merging: Aligning generated visuals, variable-length audio, and text overlays cleanly. We had to build a robust 3-tier fallback in FFmpeg (Full merge -> Video+Subtitles without audio -> Simple concat) to handle unpredictable generation edge cases without failing the user experience.
Accomplishments that we're proud of
- Successfully creating a "one-click" workflow where a simple prompt results in a fully produced, multi-modal educational video in minutes.
- Engineering a seamless React streaming interface that provides real-time progress transparency.
- Upgrading from basic storyteller cartoons to a professional "explainer video" pipeline with accurate infographics, motion graphics, and contextual chat.
- Maintaining a high standard of educational accuracy while prioritizing enjoyment and student engagement.
What we learned
We learned that the "story" is the most effective data compression tool. By wrapping facts in a narrative and reinforcing them with visual motion graphics, the AI builds strong memory anchors. We also learned how crucial "context overlap" is when prompting sequential video generation models like Veo 3 to prevent jarring visual jumps.
What's next for RAWI
- Interactive Branching: Allowing students to choose their own path in the video (e.g., "Which scientific path should we follow next?").
- SME & Curriculum Integration: Scaling the platform as a SaaS for Egyptian SMEs and schools to upload their own PDFs and automatically generate course videos.
- Multilingual Expansion: Adding regional Arabic dialects (and other languages) to ensure no student is left behind due to language barriers.
Log in or sign up for Devpost to join the conversation.