Inspiration

Over 430 million people worldwide live with disabling hearing loss. While modern captioning engines transcribe human dialogue reasonably well, they completely fail at conveying critical ambient audio: an approaching emergency siren, a screeching car tire, a blaring fire alarm, glass shattering, or someone shouting a warning. Missing these environmental acoustic cues represents a significant physical safety hazard. We built SoundSight to bridge this sensory divide by turning ambient acoustics and peripheral camera inputs into real-time, directional visual intelligence.

What it does

SoundSight operates as an ambient edge sensory copilot:

  • Real-Time Acoustic Triage: Continuously processes incoming audio streams to detect emergency events (sirens, horns, glass breaks, alarms) and sound pressure spikes.
  • Directional HUD Warnings: Maps acoustic events into high-contrast directional visual alerts on screen (e.g., “Siren Approaching from Left [85 dB]”).
  • Multimodal Contextual Synthesis: When an alert triggers, SoundSight pairs the audio signature with real-time video snapshots, running multimodal reasoning to deliver actionable heads-up advisories (e.g., "Emergency vehicle approaching on the left lane — step clear").
  • Name & Speech Wake-Detection: Transcribes spoken audio locally to highlight when the user’s name is called in noisy rooms, spotlighting the speaker.

How we built it

  • Frontend & Audio Processing: Built with a lightweight, low-latency UI using the Web Audio API and real-time audio spectrogram analysis to track frequency envelopes and decibel thresholds.
  • Multimodal AI Reasoning: Integrated Gemini Flash API with strict structured JSON output schemas to analyze incoming sound events alongside synchronized webcam frames without latency bottlenecks.
  • Edge Deployment: Configured client-side audio feature extraction with local speech recognition for immediate wake-word detection before passing alert contexts to the multimodal backend.

Challenges we ran into

  • Audio-Visual Latency Synchronization: Aligning the exact audio peak timestamp with the appropriate camera snapshot frame required careful buffering to avoid outdated context.
  • False-Positive Suppression: Tuning frequency filters so that everyday domestic noises (typing, ambient room echoes) did not trigger emergency HUD banners.
  • JSON Output Consistency: Ensuring the multimodal model strictly returned schema-compliant JSON under high-pressure event loops, which we resolved by enforcing response schemas.

Accomplishments that we're proud of

  • Delivering end-to-end incident detection and advisory generation in under 1.5 seconds.
  • Designing an intuitive, distraction-free HUD optimized specifically for high accessibility and quick situational awareness.
  • Building a functional sensory-transduction prototype that directly addresses physical safety.

What we learned

  • How to harness multimodal reasoning not just for static document Q&A, but as an active sensory translation layer for accessibility.
  • Effective stream management with the Web Audio API and real-time structured model calls.

What's next for SoundSight

  • Wearable Haptic Feedback: Synchronizing alerts with smartwatch vibration motors (directional pulse patterns).
  • Offline Edge Models: Exporting lightweight quantized audio models for standalone on-device execution without active internet access.
  • Spatial Audio Expansion: Integrating multi-microphone array direction-of-arrival (DoA) algorithms for exact 360-degree positional accuracy.

Built With

  • accessibility
  • artificial-intelligence
  • computer-vision
  • gemini-api
  • javascript
  • multimodal-ai
  • python
  • react
  • streamlit
  • web-audio-api
Share this project:

Updates

Submission history