TL;DR

Scout turns an everyday smartphone into an infrastructure-free indoor navigator for blind and low-vision users. Worn on the chest or held forward, it reads room signs, flags ground hazards (stairs, obstacles), and speaks concise clock-face guidance—no beacons, no pre-mapped building, no specialized hardware. The goal is accessibility that travels with the user: cheaper and more portable than vision systems that depend on venue infrastructure.

Inspiration

GPS dies at the lobby. For millions of people with vision loss, unfamiliar indoor spaces stay hard to navigate. Many “indoor wayfinding” products assume Bluetooth beacons, custom hardware, or pre-scanned floor plans—setup that venues often never install. Scout starts from a different premise: if you already have a phone camera, you should already have a spatial guide.

What it does

  • Sees what’s ahead — doors, people, chairs, stairs, signs, and other anchors from live camera frames
  • Reads the building — OCR for room numbers / exit text (e.g. “ROOM 204”) so destinations can be confirmed out loud
  • Flags danger early — descending stairs and path obstacles as immediateHazard with short, speakable warnings
  • Speaks in clock language — e.g. “Door at 11 o’clock, about 3.5 meters” for Person 3 TTS / guidance
  • Hands off cleanly — a verified PerceptionFrame contract so mapping (Person 2) and speech (Person 3) don’t depend on raw model JSON

How we built it

Three-layer pipeline, deliberately decoupled:

  1. Perception (Person 1) — Expo app captures compressed JPEG frames (~640×480). On Android, YOLO26n + ML Kit OCR run on-device; Gemini Flash is called only when gated (errors, low confidence, caution text). Everything is sanitized into a strict PerceptionFrame schema with clocks, distances, OCR, and hazard flags.
  2. Mapping & planning (Person 2) — Consumes that frame (anchors, obstacles, signs) to plan discrete moves (LEFT / RIGHT / STRAIGHT / STOP / ARRIVED) without owning vision.
  3. Guidance (Person 3) — Turns planner + perception speakable lines into TTS and haptics for eyes-free use.

Offline mocks, schema tests, and snapshot checkpoints let each layer integrate without a live camera or burning API quota.

Challenges we ran into

  • Latency budget — Keep capture → sanitize → speak under a usable round-trip; free-tier Gemini often needs gating, retries, and safe timeout frames (“Vision service slow…”) instead of hanging.
  • Clock-space grounding — Prompting and validation so detections land on reliable 9–3 o’clock positions Person 3 can say verbatim.
  • Phone reality vs. Expo Go — Full YOLO/OCR needs a native Android build; Expo Go is camera + cloud vision only, with quota limits.
  • Blind-first UX — Prefer audio/haptic paths over dense UI so the product stays usable without looking at the screen.

Accomplishments that we're proud of

  • Zero venue infrastructure — works from a stock phone camera
  • Strict handoff contract — PerceptionFrame + teammate helpers tested with deterministic mock snapshots
  • Hybrid perception — on-device YOLO/OCR with sparse Gemini so demand and latency stay under control
  • PRD-complete Person 1 track — Steps 1–5 + security audit + final verification, ready for Person 2/3 integration
  • Aligned with SDG 11.7 (accessible public spaces) and 10.2 (social inclusion)

What's next for Scout

  • On-device IMU fusion (accel/gyro) for dead-reckoning between vision frames
  • Stronger offline hazard models so demos don’t depend on Gemini quota
  • Spatial / binaural audio cues beyond mono TTS
  • Multi-floor flows (elevator / stair transitions, button/sign reading)
  • Finish Android native YOLO path on-device and tighten end-to-end hallway trials with Persons 2 & 3

Built With

Share this project:

Updates

Submission history