Inspiration

We were inspired by the idea of making everyday environments easier to navigate for people who may have difficulty identifying obstacles around them. We wanted to create something that could use a camera and AI to recognize objects in a person's path and communicate useful information through audio.

What it does

This led us to SeeNARIO, an AI-powered vision assistant that uses a camera to identify objects, determine whether they are on the left, center, or right of the user's view, and determine whether they could pose a potential hazard. Instead of requiring the user to look at a screen, SeeNARIO communicates important information through spoken audio.

How we built it

SeeNARIO uses a Logitech webcam connected to a Raspberry Pi to capture the user's surroundings. The captured image is sent to Google Gemini, which analyzes the scene and returns structured information about the most relevant object, its position, and whether it presents a potential danger.

The information is then converted into a short warning or description and sent to ElevenLabs, which generates human-sounding speech. The audio is played through a JBL Go speaker.

We designed Gemini's output to be simple and predictable, such as: CHAIR | CENTER | YES

SeeNARIO uses a Logitech webcam connected to a Raspberry Pi to capture the user's surroundings. We also integrated a Sensor Tile.box motion sensor, which detects movement and can trigger the webcam to analyze the surrounding environment.

Challenges we ran into

One of our biggest challenges was finding the right balance between accuracy and API availability. We found that the Gemini Flash model that performed best for our application gave us strong results for identifying objects, determining whether they were on the left, center, or right, and judging whether they were potentially dangerous. However, the model had limited usage available to us, so we had to be careful about how often we sent images for analysis and optimize our testing.

Another challenge was our original plan to use a distance sensor to determine how close an obstacle was. Distance sensors were not initially available to us during the hackathon, so we had to adapt our approach. We later obtained a Sensor Tile.box motion sensor, which allowed us to incorporate motion detection into our system. Rather than measuring distance, the sensor can detect movement and work with the webcam to determine when the surrounding environment should be analyzed.

Accomplishments that we're proud of

We are proud that we were able to turn our idea for SeeNArio into a working AI-powered vision and audio system. We successfully connected a Logitech webcam, Raspberry Pi, Gemini, ElevenLabs, and a JBL Go speaker into one pipeline.

We are especially proud of adapting our project when we ran into limitations during the hackathon. Even without access to a distance sensor and with limited usage of the Gemini model that performed best for our application, we were able to find a way to keep the core idea working. Seeing the system recognize an object, determine its position, and communicate that information through audio was a rewarding part of the project.

What we learned

This hackathon was our first time working with APIs, and we learned how to connect external AI services to a physical hardware system. Working with Gemini and ElevenLabs helped us understand how APIs can take information from one part of an application, process it, and return something that can be used by another part of the system.

We also strengthened our embedded systems skills by working with the Raspberry Pi, webcam, and speaker and learning how software and hardware interact in a real-world application.

Beyond the technical skills, we learned the importance of testing, troubleshooting, and adapting. Not everything worked exactly as we originally planned, so we had to work around limitations and make decisions that kept our core idea functional. This taught us how to approach problems more flexibly and turn an idea into a working prototype under a limited time frame.

What's next for SeeNARIO

We see SeeNARIO as a starting point for a much more capable assistive vision system. In the future, we would like to improve the accuracy and speed of object detection, add more precise ways of estimating distance, and make the system smarter about deciding which objects actually require an audio alert.

Redesign SeeNARIO as a sleek, compact, consumer-friendly assistive device. Minimize the space needed for the webcam, Raspberry Pi, and speaker by integrating the components into one small, polished enclosure. Hide exposed wiring, use smooth edges, and create a clean, modern look that feels like a finished commercial product rather than a prototype. Keep it lightweight and portable while maintaining a clear camera view and unobstructed speaker. We could implement this into headphones instead of a speaker.

Ultimately, we hope to continue developing SeeNARIO into a system that can provide useful, timely information about a person's surroundings without requiring them to constantly look at a screen.

Share this project:

Updates

Submission history