Inspiration

Recent advances in on-device computer vision made us realize how underutilized the iPhone's capabilities are, especially for accessible technology. So we decided to use these resources to improve accessibility for visually impaired individuals. We believe the future of accessibility isn't just about building better models, but also about building the systems around them.

What it does

Probe helps visually impaired users locate objects, navigate toward them, and understand their surroundings. It's designed as an assistant for anyone across the spectrum of vision loss, from mild impairment to total blindness.

How we built it

We fine-tuned a YOLO26 segmentation model on a custom dataset of household containers (water bottles, cans, and cups) to detect and localize objects. Apple's pose estimation framework tracks the user's hand in the same frame. Using segmentation instead of standard bounding boxes lets us calculate the shortest distance between the hand and the object's actual edges, rather than a rough box outline. The iPhone's LiDAR sensor then triangulates the real-world distance between hand and object; once that distance approaches zero, the object is considered found.

Users are guided to the object through visual and/or audio cues, depending on their level of vision.

Challenges we ran into

  • Finding a reliable object-detection approach. No single existing model could detect both a hand and a target object accurately enough to measure distance via LiDAR. We split the problem into two models instead. We used Apple's Pose Estimation framework for the hand, and a custom-trained YOLO26 model for the object. With more time, we'd collect a broader dataset covering many household items. However, due to the hackathon's constraints, we focused on polishing detection for one object category.
  • Getting accurate hand-to-object distance via LiDAR. Defining when an object was "found" vs. "obtained" was harder than expected. To solve this issue, we added a segmentation layer to our YOLO26 model which gave us much more precise edges and made the hand-to-object distance calculation reliable.
  • Making the app usable by people with total vision loss. Every phase of the app is voice-controlled using ElevenLabs, backed by a generative AI agent that interprets user audio and responds with relevant information. We also chose a high-contrast, colorblind-friendly palette and kept the UI deliberately simple.

Accomplishments that we're proud of

  • Our YOLO26 segmentation model reliably identifies and distinguishes between container types, including water bottles, cans, and cups.
  • We calculate the distance between a user's hand and a target object in real time by fusing the segmentation mask with LiDAR depth data.
  • Everything runs on-device.

What we learned

  • Effective methods for training a computer vision model
  • How LiDAR sensors work and how to fuse depth data with vision output
  • How to break down a hard problem, weigh different technical approaches, and make fast, pragmatic decisions on tooling under a 36-hour time limit

What's next for Probe

  • Expanding our training dataset to cover a much wider range of household objects
  • Adding full navigation features (guiding users through spaces, not just to objects)

Built With

  • apple-vision-framework
  • elevenlabs
  • swift
  • xcode
  • yolo26ml
Share this project:

Updates

Submission history