Inspiration
Around 466 million people worldwide have disabling hearing loss (WHO). They miss sounds hearing people take for granted — a doorbell, a phone, someone calling their name. And some missed sounds are dangerous: a smoke alarm at night, a siren, breaking glass, a baby crying in the next room.
I asked what AI is actually good at — not what sounds impressive. Audio event classification is one of the most mature, reliable tasks in machine learning: it doesn't hallucinate, runs fast on-device, and gives clear confidence scores. Meanwhile, existing solutions all fail somewhere: dedicated alerting hardware is expensive and uses fixed generic sound libraries; phone apps can't learn your doorbell; cloud services upload recordings of your home — a privacy non-starter for exactly the vulnerable users who need help most.
The gap was clear: a free, browser-based tool that recognizes real environmental sounds, learns your sounds, and never lets audio leave the device.
What it does
Auris turns any phone or laptop into an environmental sound awareness companion. Open the page, grant microphone permission, and it classifies sounds in real time — no install, no account, no cost.
- Priority-tiered alerts: critical sounds (smoke alarm, siren, glass breaking) trigger a full-screen red overlay + vibration — impossible to miss. High-priority sounds (doorbell, phone, baby crying) get a prominent toast + short vibration; normal sounds get a quiet toast.
- Personalized transfer learning: record your own doorbell — or your name — in 5 seconds. Auris learns to recognize your sound, not a generic one, via a 1024-dim audio embedding. Only the vector is stored, never the audio.
- Calibrated decision layer: per-class confidence thresholds (critical sounds use a low threshold — better a false alarm than a missed smoke alarm), multi-frame debouncing, cooldown, and explicit "uncertain" labeling instead of guessing.
- Real-time transparency: live spectrogram + top-5 readout — you can see what the model hears.
- Privacy-first: all inference runs on-device. No backend, no API keys, no audio upload. Works fully offline after first load (PWA).
- Demo mode: built-in test sounds feed the real inference pipeline, so anyone can experience it reproducibly.
Try it: https://georgefifth.github.io/auris/
How we built it
- TensorFlow.js + YAMNet — the model (trained on AudioSet's 521 sound classes) is bundled locally (~15MB) rather than loaded from TF Hub, enabling full offline use.
- Web Audio API + AudioWorklet — real-time capture with 16kHz resampling inside the worklet thread (main-thread fallback for older browsers).
- Inference pipeline: each 0.975s window → scores over 521 classes + 1024-dim embeddings + spectrogram; scores are mean-aggregated and softmax-normalized.
- Decision state machine: preset sounds map AudioSet classes to 9 alert categories with independent thresholds; custom sounds match by cosine similarity to stored embeddings (threshold 0.70); 3-frame debouncing + 5s cooldown prevent flicker.
- Vanilla JS + Vite — minimal, fast, static deploy on GitHub Pages with zero operating cost.
Challenges we ran into
- The model CDN broke. TF Hub started returning 403 after its migration to Kaggle. Fix: bundled the model files locally — which turned out to be a feature, enabling full offline use.
- Browser sample-rate mismatch. YAMNet expects 16kHz mono; browsers run AudioContext at 48kHz and may ignore a requested sample rate. Fix: capture at native rate and resample ourselves inside the AudioWorklet.
- Demo reproducibility. Mic quality and room noise make live demos unreliable. Fix: a demo mode that feeds decoded test samples directly into the inference pipeline — 100% reproducible.
- A confident wrong answer is worse than no answer. Tuning thresholds per class, adding debounce/cooldown, and labeling gray-zone results "uncertain" took more design thought than the model itself.
Accomplishments that we're proud of
- Zero backend, zero API keys, zero audio upload. Inference, transfer learning, and alerts all run in the browser — a privacy win and a deployment win in one.
- Personalized sound recognition commercial hardware can't do. Fixed sound libraries vs. a 5-second teach-in of your doorbell.
- It exists and it's live. Not a mockup — a working product anyone in the world can open right now, free.
- Calibrated honesty. The system knows when it doesn't know — essential for a safety tool.
What we learned
- Audio classification is mature enough to be boring — and boring is exactly what safety-critical applications need.
- On-device ML is production-practical: real-time inference in a laptop browser with no noticeable lag.
- Transfer learning doesn't require retraining — embedding similarity is real personalization with a frozen model, no GPU.
- Privacy can be a feature, not a constraint: "audio never leaves your device" is ethically important and technically simpler.
What's next for Auris
- Bluetooth vibration wearables for sleep use — a deaf person sleeping can't see a screen, but a wrist puck can wake them for a smoke alarm.
- Smart-home integration — flash lights on critical sounds (Philips Hue, Home Assistant via MQTT).
- Multiple custom sounds with incremental learning — no retraining of existing ones.
- Background-tab persistence — keep listening via Web Notifications + a persistent worker.
- Multi-language UI — currently English-only.
Prior-work disclosure (per Rule 3): Auris was originally built for UPAI-Hackdays (Sep 10, 2026) and is submitted here as clearly identified pre-existing work; the linked demo and repository are the original artifacts. AI tools assisted with drafting this submission text.
Built With
- accessibility
- deaf
- machine-learning
- on-device-inference
- privacy
- pwa
- social-impact
- tensorflow-js
- transfer-learning
- web-audio-api
Log in or sign up for Devpost to join the conversation.