Inspiration

Traditional AI tutors are text-based and slow. Lumi changes that. It's a real-time, multimodal agent that you can talk to naturally, just like a human teacher

What it does

It solves the problem of disconnected learning by letting the AI 'see' what you're working on and responding instantly with voice, creating a truly immersive educational experience.

How we built it

We used the Gemini Live API to see what you see and hear what you say in real-time, functioning as an interactive AI tutor

Challenges we ran into

To use ai.live.connect with gemini-2.5-flash-native-audio-preview-09-2025 for bidirectional streaming. To Capture microphone input (16kHz PCM) and plays back model audio using the Web Audio API.

Accomplishments that we're proud of

Building the Multimodal live agent with Gemini API

What we learned

We learnt how to build Multimodal Live Agent using the Google GenAI SDK to connect directly to the Gemini 2.5 Flash model on Google Cloud

What's next for Lumi

We work on this model to update this model with great features

Built With

  • motion
  • react
  • tailwind-css
  • vite
Share this project:

Updates