Inspiration

Finding a misplaced everyday object can be especially difficult when you cannot visually scan the room. We wanted to make that task more independent using technology people already own.

According to the American Foundation for the Blind, citing the 2024 American Community Survey, 82.1% of people with vision difficulty in the United States own smartphones. That made the ordinary smartphone camera a meaningful starting point. Rather than requiring a depth sensor or specialized wearable, we set out to combine a standard 2D camera with spoken guidance.

Seekr grew from a simple goal: help blind and low-vision people find what they need, on their terms.

What it does

Seekr is a voice-driven prototype for finding and approaching everyday objects while helping users navigate around obstacles in their path. The user connects their phone to the application, starts the camera, and says what they are looking for—for example, “a red cup.”

The system analyzes camera frames to identify the requested object, detect potential obstacles, and estimate where visible floor space could support a safe approach. It translates those observations into short spoken instructions, including turning to avoid an obstacle, taking a small step, stopping, or looking again.

The interaction continues beyond detecting the object. Seekr assesses apparent reachability, helps users adjust their route when obstacles block the way, asks the user to confirm whether they can feel the item, and accepts responses such as “yes,” “too far,” or “got it.” Users can also pause, repeat instructions, or change their target through voice commands.

The current prototype uses the phone for its camera, microphone, and audio, while a connected computer/backend coordinates processing and displays diagnostics.

How we built it

We began with camera-based object detection using Google Gemini. The initial version identified everything in view, which produced inconsistent results. We narrowed detection to the requested object and refined the model’s structured output, giving the rest of the system a more reliable target to search for.

We added ElevenLabs speech output and connected the phone camera through a QR-code pairing flow, moving the system from a desktop demonstration to a portable search tool. This made the guidance accessible to users while allowing the backend to receive live visual information from the phone.

We then added NVIDIA’s SegFormer semantic segmentation model to identify likely floor regions. The latest navigation branch calculates routes from segmentation masks rather than relying on Gemini’s broad scene assessments, helping the system choose safer movement directions and reducing dependence on general scene descriptions.

We built a navigation controller with stages for searching, aligning, approaching, recovering a lost target, and confirming pickup. Timestamps, phone-orientation checks, and session tracking prevent delayed results from triggering new movement instructions, allowing the components to work together as a coordinated and responsive control loop.

We simplified the phone interface and made ElevenLabs Scribe the primary speech-to-text service. The final stack combines a Python/FastAPI backend, a JavaScript browser interface, WebSocket phone communication, Gemini scene analysis, SegFormer segmentation, and ElevenLabs speech input and output. Each part contributes a distinct function: the interface supports interaction, the communication layer connects the phone and backend, the vision models interpret the environment, and the speech services provide hands-free input and feedback.

We also explored depth estimation and bird’s-eye mapping. These experiments informed our understanding of the navigation problem, while the submitted branch focuses on guidance from the current camera view. Although they are not part of the final system, they helped evaluate possible future improvements to spatial reasoning and could support more complete navigation in later versions.

Challenges we ran into

Getting useful object detection. Our initial approach asked the model to describe too much at once. We improved the implementation by focusing on the requested target, explicitly allowing “not visible,” and refining the information needed for guidance.

Keeping instructions connected to the present. Camera frames, model responses, and speech all arrive at different times. An observation could become outdated while the user turned or changed targets. We added checks for duplicate frames, session changes, and significant rotation, along with requirements for fresh observations before further movement cues.

Preventing the application from overwhelming the user. Repeated observations could produce repeated speech, while slow requests could accumulate unnecessary work. We introduced duplicate-cue suppression, limits on simultaneous requests, and cancellation of results that belonged to an earlier interaction.

Choosing a manageable navigation approach. We explored richer depth and room-mapping systems, but they introduced additional calibration, tracking, and integration challenges. We narrowed the main prototype to object finding and routes through the current camera image, giving us a more focused system to build and inspect.

Accomplishments that we're proud of

We are proud of integrating automatic speech recognition, natural-language intent parsing, computer-vision-based scene understanding, and interactive visual search into a cohesive multimodal pipeline. Spoken requests are transcribed and mapped to structured search intents, relevant visual content is retrieved and analyzed using object detection and scene segmentation, and the resulting context is used to generate adaptive guidance. Users can provide corrective feedback or confirmation through the interaction loop, enabling incremental intent refinement and robust task completion.

The most significant achievement is the coordination between these components. Gemini identifies and describes the target, SegFormer supplies image segmentation, our controller decides which instruction is appropriate, and ElevenLabs supports spoken interaction. This integration is significant because it transforms several specialized tools into a unified system that can perceive its surroundings, interpret visual information, make context-aware decisions, and communicate them naturally. Rather than relying on a single model to perform every task, the system combines the strengths of each component, improving flexibility, reliability, and accessibility while creating a foundation for real-time assistance in practical environments.

We also kept accessibility central to the design. The phone interface is deliberately minimal, and the core interaction uses speech rather than requiring the user to interpret bounding boxes or a visual map. A separate desktop dashboard lets us inspect the system without placing that complexity in the user’s interface.

All of this begins with a standard smartphone camera, supporting our goal of reducing the need for specialized hardware.

What we learned

Nearly every major part of Seekr was new to us. We learned how to connect a phone camera to a backend, work with multimodal models, interpret segmentation masks, build route-selection logic, and coordinate speech recognition with audio playback.

We learned that recognizing an object is only one part of helping someone retrieve it. The system also needs to account for changing camera views, uncertain observations, obstacles, timing, and the difference between seeing an object and being able to reach it.

We also learned to separate what each technology can establish. A bounding box identifies an image location; segmentation classifies pixels; neither automatically provides reliable physical distance or body clearance.

Most importantly, we learned that accessibility involves the entire interaction. Instruction timing, repetition, recovery, and user confirmation matter alongside model accuracy. Automated checks help us examine those behaviors, but testing with blind and low-vision users is essential to understanding whether the experience is useful.

What's next for Seekr

Our next priority is partnering with blind and low-vision users to evaluate and improve the current object-finding experience. We will focus on successful retrievals, clear and accurate instructions, timely responses, and how effectively the system communicates uncertainty and supports users in recovering when needed.

We also want to:

  • Improve distance and reachability estimates, including the distinction between visible floor space and a passage with enough physical clearance.
  • Maintain targets more reliably through turns, temporary occlusion, and scenes containing several similar objects.
  • Handle more complex environments, including moving obstacles and hazards that floor segmentation alone cannot establish.
  • Strengthen reliability, including browser compatibility, network interruptions, speech behavior, and regression testing.

Longer term, we want to expand from finding individual objects toward broader indoor assistance while preserving the accessibility of an ordinary smartphone camera.

Built With

Share this project:

Updates

Submission history