🌟 Inspiration

Our inspiration came from wanting to make everyday navigation more accessible for people with vision difficulties. Tools like canes and guide dogs provide invaluable assistance, but we wanted to explore how modern computer vision and voice technology could provide an additional layer of awareness about the world around a person.

🤔 What it does

VisionCompanion helps people with vision difficulties navigate their daily lives with greater independence. Where a cane may leave gaps, VisionCompanion provides an additional layer of awareness by using Gemini’s computer vision to identify important objects and obstacles in the user’s surroundings. [^1] It then communicates these insights through a voice companion powered by ElevenLabs Voice Agents, helping users better understand and navigate the world around them.

[^1]: VisionCompanion is a complementary application, not a replacement for traditional assistance tools such as blind canes.


🥞 Our TechStack

Our front end was written in TypeScript using React (Vite). Image processing is done through FastAPI calls to Gemini 3.5 Flash Lite, and callouts are voiced by ElevenLab's voice agents.

🚧 Challenges we ran into

  • Initial API calls
  • High token usage
  • Designing the front-end
  • Setting up the website

🎖️ Accomplishments that we're proud of

Most of us had our first experience with AI vision detection, and doing the research on different tools we could use helped us learn a lot about it. We added a few accessibility features (large text, English/French UI, voice accents, etc.) which would've not been possible if there was no time.

🔮 What's next for VisionCompanion

  • Support for a larger pool of languages
  • Fine tuning of vision (potential hybrid of API calls and in-browser learning)
  • Haptics (vibrations etc.)

Built With

Share this project:

Updates

Submission history