Inspiration
I'm big into health tech, and lately I've been paying attention to how fast AR glasses and AI cameras are blowing up, stuff like Meta's smart glasses. And it got me thinking, all this tech is being built just to make everyday life a little more convenient for people who honestly don't need it that badly. Meanwhile blind people are still walking around with a cane, going slow, tapping the ground just hoping they don't trip or walk into something. That felt backwards to me. So I wanted to take that same camera and AI tech and actually put it toward something that matters, helping blind people move through the world normally, at a normal pace, without having to be so cautious about every step they take.
What it does
guideGlass is camera glasses plus a web app. The glasses stream what the wearer sees to an AI model, which speaks out hazards in front of them in real time, like a wall or a clear path. The web app also gives turn by turn walking directions to a typed destination using the phone's GPS, spoken out loud as they walk.
How we built it
An ESP32-CAM sends frames over MQTT to HiveMQ. A Python backend picks up the latest frame every second or two and sends it to Gemini, which returns a short spoken warning. That gets pushed to the browser over a WebSocket and read aloud with the Web Speech API. The browser also pulls phone GPS and calls the Google Maps Directions API for turn by turn navigation, layered on top with its own voice.
Challenges we ran into
Keeping the frame stream from flooding the backend, so I just keep the latest frame and drop the rest. Gemini's free tier daily limit also cut me off mid demo since I was calling it too often. Getting two voices, hazard warnings and turn instructions, to not talk over each other took some thought too.
Accomplishments that we're proud of
Getting the whole pipeline working end to end in the time I had, real camera, real AI model, real spoken guidance. Adding full navigation on top with no new hardware, just the phone already running the app.
What we learned
A lot of it comes down to plumbing, not the AI itself: throttling frames, handling reconnects, keeping the two voices from colliding. Also learned how to balance real time feel against free tier rate limits.
What's next for GuideGlass
A real speaker and mic on the glasses so it doesn't need a phone screen, better obstacle detection (telling a person from a car from a wall), and maybe onboard depth sensing for when there's no strong connection.
Built With
- api
- based-on-what's-actually-in-the-project:-python
- css
- esp32-cam
- fastapi
- gemini
- google-maps-directions-api
- gps
- hivemq
- html
- javascript
- mqtt
- web-speech-api
- websocket
Log in or sign up for Devpost to join the conversation.