🌟 Inspiration
Our inspiration came from wanting to make everyday navigation more accessible for people with vision difficulties. Tools like canes and guide dogs provide invaluable assistance, but we wanted to explore how modern computer vision and voice technology could provide an additional layer of awareness about the world around a person.
🤔 What it does
VisionCompanion helps people with vision difficulties navigate their daily lives with greater independence. Where a cane may leave gaps, VisionCompanion provides an additional layer of awareness by using Gemini’s computer vision to identify important objects and obstacles in the user’s surroundings. [^1] It then communicates these insights through a voice companion powered by ElevenLabs Voice Agents, helping users better understand and navigate the world around them.
[^1]: VisionCompanion is a complementary application, not a replacement for traditional assistance tools such as blind canes.
🥞 Our TechStack
Our front end was written in TypeScript using React (Vite). Image processing is done through FastAPI calls to Gemini 3.5 Flash Lite, and callouts are voiced by ElevenLab's voice agents.
🚧 Challenges we ran into
- Initial API calls
- High token usage
- Designing the front-end
- Setting up the website
🎖️ Accomplishments that we're proud of
Most of us had our first experience with AI vision detection, and doing the research on different tools we could use helped us learn a lot about it. We added a few accessibility features (large text, English/French UI, voice accents, etc.) which would've not been possible if there was no time.
🔮 What's next for VisionCompanion
- Support for a larger pool of languages
- Fine tuning of vision (potential hybrid of API calls and in-browser learning)
- Haptics (vibrations etc.)
Built With
- css3
- elevenlabs
- fastapi
- gemini-api
- godaddy
- react
- typescript
- vercel
- vite
Log in or sign up for Devpost to join the conversation.