Inspiration
2.2 billion people globally live with a vision impairment. But, cities aren't built with them in mind. Tasks most of us never think about, dodging an obstacle, recognizing a friend, become daily obstacles.
Recently, there's also also a new threat: AI voices. It takes just 3 seconds of audio to clone a voice, and people spot a fake only 55% of the time (barely better than a coin flip!). For someone who relies on their ears to know who they're talking to, that can pose as a real safety risk.
This hackathon, we aim to answer this question: what if a familiar pair of glasses could see on your behalf, and tell you what matters, the moment it matters?
What it does
See Through is a wearable companion that clips onto everyday glasses and turns what the camera sees into guidance, so recognition doesn't depend on a voice that could be faked.
- Recognizes loved ones: enroll familiar faces, and the glasses confirm who's in front of you out loud.
- Spots hazards: flags people, cars, and stop signs, with direction and rough distance ("car, ahead-right, ~4m").
- Talks back: fully hands-free, push-to-talk voice control powered by xAI to check profiles, ask "Where am I?", or "Repeat the last alert."
- Knows your safe spaces: voice-managed GPS places like Home, with status and arrival awareness.
- Understands context: Snowflake Cortex turns raw detections into natural spoken cues, daily-pattern memory, and family alerts.
- Protects privacy: camera photos never touches the AI.
The hardware
See Through runs on a custom rig we assembled and fitted onto a pair of glasses:
- Sense board: Seeed XIAO ESP32-S3 Sense, a fingernail-sized module with dual wireless (Wi-Fi + Bluetooth 5 LE) doing all the on-glasses work.
- Camera: an OV2640 that fires short three-photo JPEG bursts over Bluetooth LE instead of streaming video, keeping power draw and bandwidth low.
- Antenna: external 2.4 GHz patch antenna that supports 300ft+ BLE link to the paired laptop.
- Battery: a LiPo pack with hand-soldered leads so the glasses run untethered.
P.S.: We also built an interactive 3D model of the rig into the app so anyone can rotate it and tap each component to learn what it does.
How we built it
The hard part was getting reliable output from flaky, low-power hardware. A laptop bridge speaks BLE to custom OpenGlass firmware, reassembles the chunked JPEG payloads, and drops any frame that decodes as truncated or corrupt.
Every ~3s we fire a 3-shot burst, score each frame on sharpness and face presence, and forward only the best one. Below goes into depth the equations that we used:
Sharpness of grayscale frame \( variant-of-Laplcian sharpness score \):
$$ S = \operatorname{Var}(\nabla^2 I) $$
We keep the sharpest frame only if it clears a floor tuned to the real burst range (~5–40):
$$ S_{\max} \ge 15 \Rightarrow \text{keep}, \quad \text{else } \texttt{retake_needed} $$
Face gate (pass/fail, not blended into \( S \)): InsightFace must return a valid embedding on the kept frame, otherwise the scene is marked no_face_detected.
Only a frame passing both gates advances to InsightFace (face recognition) and YOLO (obstacle detection). This filters out the motion blur and misframes you get from a shaky, wide-angle wearable camera.
Recognition runs locally, InsightFace for faces, YOLO for obstacles (with bearing and monocular distance), so it works fully offline. With Wi-Fi, Snowflake Cortex handles scene reasoning, narration, and family alerts; xAI drives transcription and function-calling voice control. Only structured metadata leaves the laptop (JPEGs never do).
Result: real-time speech and alerts on glasses that stay lightweight, low-power, and private by default.

Log in or sign up for Devpost to join the conversation.