Inspiration
Think about walking into a room you've never been in, without being able to see it. Where's the door? Where's the table? Most blind people learn a new place by having someone walk them through it, on the day. I wanted them to be able to learn it the night before, at home, by sound.
I was first working on a tool for hearing how a speaker would sound in your own room. However, on day four I saw that its two main pieces, a 3D room built from a phone video and sound placed inside it, could do something that matters more.
There is research behind the idea: in a 2012 study, blind people who learned a building's layout from an audio game could then find their way through the real building.
What it does
- Film. A sighted helper films one slow lap of a room on a phone.
- Check. The laptop turns the video into a 3D room and lists what it found: doors, tables, chairs, stairs. The helper renames, removes or adds objects, because a wrong label would mislead someone who can't see it.
- Explore. The explorer puts on headphones and walks the room with the keyboard:
- each object says its name from where it really is;
- Space sweeps the room clockwise ("TV, table, door");
- Tab picks a place, and Enter leads there with a pulse that speeds up as you get closer;
- walking into something plays a thud and says what it was.
How I built it
- Video to 3D. ffmpeg cuts the video into frames, COLMAP works out where the camera was for each one, and OpenSplat trains a 3D Gaussian splat on the GPU. Each step runs in a Docker container.
- Finding objects in 3D. OWL-ViT finds 16 kinds of object in 24 video frames. Each 2D box is lifted into the room using the camera positions, and sightings of the same object from different frames are merged.
- Measuring the room. The floor, walls and ceiling are fitted to the 3D points, and the scale comes from the height the phone was held at. This gives the room's size in metres and stops the explorer walking through walls.
- 3D audio in the browser. The Web Audio API's HRTF panning places each voice where the object is, and the room's echo is computed from its size.
- The app. Next.js, React and three.js, with Spark to draw the splat, and a Node.js server that queues builds and reports progress.
Everything runs on one laptop with free, open-source tools. Nothing goes to the cloud. I built it solo with Claude Code + other AI tools as pair programmers.
Challenges I ran into
- Slow builds. A build of a 43-second video took 10 min 26 s. Moving COLMAP's database off the Windows folder mount and onto a Docker volume cut it to 4 min 41 s.
- A classroom of identical chairs. COLMAP placed only 2 of 199 frames, because its first guess at the lens was far too narrow for phone video. Changing the guess to match a phone's wide lens placed all 199.
- The Docker image build froze the laptop until I capped the number of parallel compile jobs.
Accomplishments that I'm proud of
- A phone video becomes a room you can walk by ear in under 5 minutes, on a laptop.
- Walking into the door says "door", not "wall".
- 282 automated tests cover walking, collisions, spoken directions and object placement, and all pass.
What I learned
- A missed door is safer than a wrong one. That is why a person checks the list before anyone explores.
- Directions have to work without a screen. Clock positions and rough distances ("door, 6 o'clock, about 4 metres") were easier to follow than angles.
- 3D reconstruction depends on its first guesses. One wrong lens setting was the difference between a room and nothing.
What's next for Hearify
Hearify hasn't been tested with blind users yet, so it makes no claim to help until it has. That test is the next step, with an orientation and mobility instructor. After that come whole buildings: a school, a clinic, a first day at work.
Log in or sign up for Devpost to join the conversation.