Inspiration
Over 430 million people worldwide live with disabling hearing loss. While modern captioning engines transcribe human dialogue reasonably well, they completely fail at conveying critical ambient audio: an approaching emergency siren, a screeching car tire, a blaring fire alarm, glass shattering, or someone shouting a warning. Missing these environmental acoustic cues represents a significant physical safety hazard. We built SoundSight to bridge this sensory divide by turning ambient acoustics and peripheral camera inputs into real-time, directional visual intelligence.
What it does
SoundSight operates as an ambient edge sensory copilot:
- Real-Time Acoustic Triage: Continuously processes incoming audio streams to detect emergency events (sirens, horns, glass breaks, alarms) and sound pressure spikes.
- Directional HUD Warnings: Maps acoustic events into high-contrast directional visual alerts on screen (e.g., “Siren Approaching from Left [85 dB]”).
- Multimodal Contextual Synthesis: When an alert triggers, SoundSight pairs the audio signature with real-time video snapshots, running multimodal reasoning to deliver actionable heads-up advisories (e.g., "Emergency vehicle approaching on the left lane — step clear").
- Name & Speech Wake-Detection: Transcribes spoken audio locally to highlight when the user’s name is called in noisy rooms, spotlighting the speaker.
How we built it
- Frontend & Audio Processing: Built with a lightweight, low-latency UI using the Web Audio API and real-time audio spectrogram analysis to track frequency envelopes and decibel thresholds.
- Multimodal AI Reasoning: Integrated Gemini Flash API with strict structured JSON output schemas to analyze incoming sound events alongside synchronized webcam frames without latency bottlenecks.
- Edge Deployment: Configured client-side audio feature extraction with local speech recognition for immediate wake-word detection before passing alert contexts to the multimodal backend.
Challenges we ran into
- Audio-Visual Latency Synchronization: Aligning the exact audio peak timestamp with the appropriate camera snapshot frame required careful buffering to avoid outdated context.
- False-Positive Suppression: Tuning frequency filters so that everyday domestic noises (typing, ambient room echoes) did not trigger emergency HUD banners.
- JSON Output Consistency: Ensuring the multimodal model strictly returned schema-compliant JSON under high-pressure event loops, which we resolved by enforcing response schemas.
Accomplishments that we're proud of
- Delivering end-to-end incident detection and advisory generation in under 1.5 seconds.
- Designing an intuitive, distraction-free HUD optimized specifically for high accessibility and quick situational awareness.
- Building a functional sensory-transduction prototype that directly addresses physical safety.
What we learned
- How to harness multimodal reasoning not just for static document Q&A, but as an active sensory translation layer for accessibility.
- Effective stream management with the Web Audio API and real-time structured model calls.
What's next for SoundSight
- Wearable Haptic Feedback: Synchronizing alerts with smartwatch vibration motors (directional pulse patterns).
- Offline Edge Models: Exporting lightweight quantized audio models for standalone on-device execution without active internet access.
- Spatial Audio Expansion: Integrating multi-microphone array direction-of-arrival (DoA) algorithms for exact 360-degree positional accuracy.
Built With
- accessibility
- artificial-intelligence
- computer-vision
- gemini-api
- javascript
- multimodal-ai
- python
- react
- streamlit
- web-audio-api
Log in or sign up for Devpost to join the conversation.