Inspiration
The inspiration for Heard came from the common frustration with modern AI productivity: voice dictation is undeniably faster and more effortless than typing, but it breaks down completely in public. If you are on a packed bus, in a quiet library, or surrounded by strangers, starting a loud voice prompt feels intrusive and awkward. We realized there was no middle ground between slow, physical typing and talking out loud.
That led to an idea: What if you could articulate mouth shapes, and have AI read your lips in real time? This would give you the benefits of dictation without any of the awkwardness
As we worked on the concept, we unlocked its true impact: Universal accessibility. What started as a private dictation hack evolved into a liberating communication bridge for individuals with speech, vocal, or hearing differences. By streaming the recognized text into ElevenLabs’ API for instant text-to-speech, and pairing it with a real-time, speaker-diarized transcript panel for incoming audio, we created a complete two-way system—allowing non-verbal and late-deafened individuals to participate naturally in dynamic, multi-person conversations without relying on manual typing or specialized hardware.
What it does
Heard is a real-time, bidirectional accessibility platform that bridges the gap between silent articulation and spoken multi-party conversation:
Silent Lip-Reading to Speech (Outgoing): Accesses your front camera to track mouth shapes, lip geometry, and facial landmarks in real time as you silently mouth words. It translates those silent movements into text and streams audio out via the ElevenLabs TTS API so surrounding listeners can hear you clearly.
Live Speaker-Diarized Chat Panel (Incoming): Captures multi-person audio through the microphone, transcribes speech live, and isolates unique speaker voices. It renders incoming conversation line-by-line in a side panel next to the camera feed—functioning like a Twitch chat—so deaf or hard-of-hearing users can identify who is speaking to them in real time.
Together, these features create a seamless, low-latency loop where you can silently mouth your thoughts to be spoken aloud, while simultaneously reading an organized, color-coded stream of everyone talking back to you.
How we built it
Lip reading (ML): We fine-tuned Auto-AVSR (LRS3_V_WER19.1), a visual speech recognition model, and exported it to ONNX. We quantized it to int8, which shrinks it from 775 MB to 203 MB, so it runs fully in the browser. Preprocessing matches training exactly: MediaPipe face detection finds the eyes, nose and mouth, and the frames are smoothed and warped to a mean face. We then crop an 88×88 grayscale mouth region. A CTC decoder turns the model output into text.
Two modes: Speed runs the model in-browser with ONNX Runtime Web (WASM, optional WebGPU), so nothing leaves your device. Accuracy sends mouth crops to a FastAPI endpoint running beam search with a language model on a RunPod GPU. It falls back to speed mode if the server is unavailable.
Frontend: React 19, TypeScript, Vite, Tailwind CSS and shadcn/ui. MediaPipe tasks-vision handles face tracking on the webcam feed. The side panel shows the diarized transcript, with each speaker color-coded.
Backend: Python, FastAPI and Uvicorn, containerized with Docker. WebSockets stream data between the browser and server. The server sends recognized text to ElevenLabs for text-to-speech. It also transcribes and diarizes incoming audio for the chat panel. API keys stay on the server and never reach the browser bundle.
Challenges we ran into
Lip Synchronization & VSR Accuracy: Achieving low latency while maintaining an acceptable margin of error was extremely difficult. Visual Speech Recognition (VSR) inherently suffers from "homophenes" (words that look visually identical on the lips, like "pack" vs. "back"). Balancing speed with low error rates required extensive time consuming fine-tuning and context-filtering to ensure smooth, natural text-to-speech output.
Real-Time Audio Diarization in Noisy Environments: Speaker diarization proved far more complex than expected. Disentangling distinct voice signatures and accurately mapping them to labeled speakers becomes exponentially harder when background noise, overlapping speech, or room echo are introduced into the microphone stream.
Data Collection & Time Constraints: Training and fine-tuning our VSR models under hackathon deadlines was a massive hurdle. Because specialized visual speech datasets are limited, we had to spend hours recording custom datasets of ourselves articulating specific sentence structures to establish ground-truth reference points for our landmark tracking pipeline.
Accomplishments that we're proud of
working deployed prototype that accomplishes what we set out to do
What we learned
We gained hands-on experience with model quantization, converting weights to lighter formats to run real-time visual speech models directly in the browser. We also got exposed to low-latency streaming architectures by leveraging WebSockets and asynchronous pipelines to process computer vision and audio streams concurrently.
Built With
- blazeface
- docker
- fastapi
- machine-learning
- mediapipe
- onnx
- python
- react
- shadcn
- tailwind
- tidb
- typescript
- uvicorn
- vite
- webgpu
Log in or sign up for Devpost to join the conversation.