Inspiration
Traditional AI tutors are text-based and slow. Lumi changes that. It's a real-time, multimodal agent that you can talk to naturally, just like a human teacher
What it does
It solves the problem of disconnected learning by letting the AI 'see' what you're working on and responding instantly with voice, creating a truly immersive educational experience.
How we built it
We used the Gemini Live API to see what you see and hear what you say in real-time, functioning as an interactive AI tutor
Challenges we ran into
To use ai.live.connect with gemini-2.5-flash-native-audio-preview-09-2025 for bidirectional streaming. To Capture microphone input (16kHz PCM) and plays back model audio using the Web Audio API.
Accomplishments that we're proud of
Building the Multimodal live agent with Gemini API
What we learned
We learnt how to build Multimodal Live Agent using the Google GenAI SDK to connect directly to the Gemini 2.5 Flash model on Google Cloud
What's next for Lumi
We work on this model to update this model with great features
Built With
- motion
- react
- tailwind-css
- vite


Log in or sign up for Devpost to join the conversation.