CaneOS
Inspiration
A white cane has been keeping blind people safe for almost a century, and it still can't see a branch at head height, an open cabinet door, or a curb two steps ahead. It only knows what it physically touches. 🦯
Meanwhile, Ray-Ban Meta glasses already do the "AI describes what you're looking at" thing, but you have to remember to ask, then wait for an answer. That's great for curiosity, not so great when something's about to hit you in the face. We wanted the opposite: a system that's always watching and speaks up the second it matters, clipped onto a tool people already trust, not a new $300 gadget to learn.
What it does
Think of it as two brains working at two different speeds. ⚡ One reacts instantly, no thinking required, like a knee-jerk reflex. 🧠 The other actually understands what's around you and explains it, like a friend narrating the room.
- ⚡ Instant reflexes: 3 sensors (left, right, up) catch what a cane can't — branches, chest-height obstacles, fast hazards. Closest one wins, watch buzzes, done. No spam, no confusion.
- 🧠 Real understanding: a camera + Gemini figure out what the hazard actually is and say it out loud through your AirPods. Not just "beep," but "low branch ahead, duck!"
- 💬 On-demand awareness: ask about your surroundings anytime, and it remembers what it already told you.
- 🆘 Safety net: big hazards or a manual SOS button pull your location and alert your emergency contacts instantly.
- ⚙️ Fully accessible app: adjustable haptics, narration on/off, live location sharing, login, and a history of past incidents.
How we built it
- 🔧 Hardware: Arduino UNO Q + 3 ToF sensors + an OAK-1-AF camera that runs on-device object detection, zero network delay.
- ⚡ Reflex path: polls sensors ~15x/sec, figures out the closest hazard, buzzes the Watch in a smart, non-spammy pattern.
- 🧠 Cognition path: enriches camera detections with real distance, sends it to Gemini for analysis, avoids repeating itself, then speaks the result through ElevenLabs.
- 🩺 Health check: a heartbeat from the camera lets us catch failures (loud or silent) and gracefully fall back to haptics-only instead of going dark.
- 📱 Swift app: three WebSocket channels feed ElevenLabs audio, Watch haptics, Backboard memory, MongoDB Atlas (contacts + history), and the SOS/geolocation flow.
Challenges we ran into
- 🔗 Contract drift: splitting hardware and software across teammates meant agreeing on exact data shapes before building, not after. We caught a WebSocket mismatch and a missing direction value this way, both would've silently broken things.
- Network issues: trouble connecting between the various software and hardware programs.
Accomplishments that we're proud of
The two-brain architecture actually held up under real stress-testing, not just in theory. We killed the API key, simulated a camera crash, and flooded the sensors with chaos, and every time, the system degraded gracefully instead of breaking. Camera dies? You still get haptics. That resilience is what we're most proud of. 🙌
What we learned
Build for failure from day one, not as an afterthought, especially when the whole point is keeping someone safe. We also learned how to collaborate across a real language barrier (Python backend, Swift frontend, two teammates with no Mac!) by treating every handoff like a contract, not a guess.
What's next for CaneOS 🚀
- Swap mocked sensors/camera for the real hardware, end to end.
- Explore stereo depth for richer spatial awareness.
- A carefully tuned fall-detection heuristic on the Watch.
- Real-world testing with blind and low-vision users.
- Keep chasing the mission: proactive safety without an expensive new gadget.
Built With
- arduino-uno-q
- backboard.io
- elevenlabs
- geminiapi
- javascript
- json
- mongodb
- oak-1-af
- python
- resend
- swift
- tof-sensors
- vercel
Log in or sign up for Devpost to join the conversation.