Inspiration

Navigating a noisy room is isolating when you can't hear it. The signals everyone else absorbs for free — the doorbell, a knock, a siren, glass breaking — arrive as nothing at all if you're d/Deaf or hard of hearing. Phone transcription apps help with conversation, but they listen from a microphone in your hand and show an unattributed wall of text: you get the words, not the who or the where. And nothing in them tells you that a sound needs you now.

We built C-HUD Vision (Cap Heads-Up Display) because the missing sense is spatial: not just what happened, but which way it came from and whether it deserves your attention. A phone has one microphone — it can detect a sound, it cannot locate one. Four microphones on a cap can, and the cap is already on your head, which is the frame that matters.

What it does

C-HUD Vision turns an ordinary cap into a real-time situational-awareness device, using the phone you already own as the display:

  • Places the sound. A bearing marker on the phone's camera view for anything the camera can see; the cap's own four-mic quadrant — left / right / front / back — for anything it can't. Direction is reported with an error bar (accuracy_deg), and a straight line of microphones is front/back ambiguous, so an ambiguous event is drawn as two mirrored candidates rather than one confident guess.
  • Names it. YAMNet (521 AudioSet classes) over a 0.975 s window labels the event — doorbell, knock, alarm, siren, glass, car horn, dog, speech — with its confidence shown on screen.
  • Ranks it. Urgency tiers: an alarm, siren or smoke detector takes the whole screen and suppresses everything else. Everything below that is yours to filter, with important and quiet modes.
  • Finds the voice. Speech events get a caption bubble anchored to the face that produced them, and a speaker playing a recording gets a marker with no face anchor, labelled as playback.
  • Keeps the wearer in the room. No headset, no AR glasses, nothing new to buy: the cap is the sensor, the phone is the screen, and it stays in your pocket until it has something to say.

How we built it

  • Hardware. Four Adafruit ICS-43434 I2S microphones on an ESP32-S3, as two stereo pairs on two sample-locked I2S buses. The firmware runs the direction-finding on the device itself — per-mic RMS, an imbalance heuristic across the left/right and front/back pairs, and an active gate so the microphones' own noise floor can't chase noise — and broadcasts a small hat_status JSON packet over UDP (and USB serial) at ~10 Hz instead of streaming raw audio.
  • Backend pipeline (Python, on a laptop in the room). Ingest from UDP / USB / a WAV file / the HUD's own microphone into per-channel ring buffers → onset detection → direction of arrival (GCC-PHAT per pair, SRP-PHAT once three or more microphones are live, with accuracy_deg derived from the measured delay spread, not from a confidence score) → YAMNet TFLite classification → faster-whisper on a worker thread → fusion with camera observations → urgency tiers → a FastAPI WebSocket the HUD subscribes to. The classifier is fed a delay-and-sum beam steered to the current bearing, never a raw four-channel sum.
  • The camera is a sensor too. The HUD already owns the webcam, so it forwards face boxes and mouth activity over the same socket; the backend turns a face into a hat-frame bearing and uses it to break the linear array's front/back tie. A calibration routine measures the array's effective spacing against the camera while someone talks, instead of trusting a CAD drawing.
  • HUD (Vite + TypeScript). getUserMedia camera with a canvas 2D overlay: bearing compass, per-event markers with class, confidence and accuracy, edge chevrons when a bearing falls outside the field of view, mirrored pairs for ambiguous events, urgency tiers, and caption bubbles pinned to faces via MediaPipe face landmarks (with a pose-based fallback for distant speakers). Positions interpolate instead of snapping, markers age and fade, the phone's orientation sensor corrects head yaw, and the phone can stream its own microphone to the backend so it becomes the sensor.
  • Nothing leaves the room. No cloud service is in the loop: the model file and class map are committed to the repo, classification and speech-to-text run locally, and the only network traffic is the cap's UDP broadcast and a WebSocket on the same machine.

Challenges we ran into

  • A straight line of microphones cannot tell front from back. That's a property of the geometry, not a bug, so we made the UI say so: ambiguous: true renders both mirrored candidates, and the camera resolves the tie only when a face is actually there. Pretending otherwise would have been the fastest way to lose a technical judge.
  • Honest error bars. accuracy_deg is mandatory on every event and comes from the measured spread of the delay estimate. When nothing can localize a sound, the event still arrives with its class but flagged unlocalized, and the marker fades out instead of pointing somewhere false.
  • The laptop's own microphone pair is not an array. Its two capsules have no inter-channel baseline (under 5 mm by three independent methods), so the coherence gate correctly refuses to report an angle on it — which is why the HUD's microphone is the default sensor and the real array is the cap.
  • Latency vs. accuracy. Balancing a 0.975 s classification window and speech-to-text against a display that must feel instant. We ended up measuring every stage separately — onset to backend, backend to client, round-trip ping — because a single end-to-end number hides which stage regressed.
  • Overload is the failure mode. An awareness device that shows everything becomes noise. Urgency tiers and mode filtering are not polish; they're the feature (a finding we owe to SoundWatch).
  • Networks at the venue. Client-isolated WiFi silently swallows the cap's UDP broadcast, and the phone needs a secure context for camera access — both of which look like "the hardware is broken".

Accomplishments that we're proud of

  • Measured, not asserted. Onset to screen at p50 378 ms; YAMNet inference at 5.65 ms per 0.975 s window; direction on synthetic audio at ±60° in 30° steps to ±1.7° on a laptop pair and ±6.0° on the modelled hat geometry; 150/150 packets with zero sequence gaps. Every number in the repo ships with the command that produced it.
  • A real wearable on real hardware. Four I2S microphones on two sample-locked buses, direction-finding on the ESP32 itself, and a phone HUD that needs no headset, no glasses and no extra hardware.
  • A system that refuses to lie. Mirrored candidates for ambiguity, mandatory error bars, and a source: none path that reports the class while admitting it doesn't know where the sound is.
  • Diagnostics that earn their place. The HUD's status chip (model hash, git revision, transport, per-mic health, noise floor, frame rate, ping latency) is what turned "why is nothing showing up" from an hour of guessing into a five-second read.

What we learned

  • Direction is the missing sense, and uncertainty is part of the answer. A one-dimensional array is ambiguous by construction, so the interface has to carry the doubt — that was the single most useful design decision we made.
  • Local-first is enough. 521-class sound classification and speech-to-text both run on a laptop in the room. There is no cloud dependency to explain, and no privacy story to hedge.
  • For accessibility, filtering beats volume. An alarm that outranks everything is worth more than a stream of everything.
  • Measuring the pipeline is the pipeline. Two of our worst hours came from debugging "bad audio" that was actually a bad model file — and one false "broken microphone" diagnosis that a stimulus-verification check now prevents.

What's next for C-HUD Vision

  • A measured microphone bar. The TDOA path is implemented and self-test verified; with a machined bar that has a real, measured baseline it runs live and replaces today's coarse quadrant with a true bearing.
  • The LED strip on the brim. Screen-free direction — the output that makes this a wearable rather than a phone accessory, and the piece we designed but did not get wired.
  • Elevation from a crown microphone, turning a bearing into a direction in three dimensions.
  • A vibrotactile band along the cap for silent, eyes-free cues.
  • A flexible PCB in the lining, and per-room calibration so the wearer never re-measures anything.

Built With

Share this project:

Updates

Submission history