Inspiration
A stroke survivor can recover movement, relearn balance, rebuild a life — and still be unable to say "I'm in pain" or "call my daughter." 22–58% of acute stroke patients develop dysarthria; 80–95% of people with ALS lose natural speech entirely; ~90% of Parkinson's patients develop it; 43% of children with cerebral palsy have impaired or no speech.
The tools that exist assume three things these patients often lack: expensive hardware ($5,000–15,000 speech-generating devices), reliable internet for cloud AI, and English. The largest unmet population — elderly stroke survivors and neurodegenerative patients in low- and middle-income countries — has none of the three. We built Svara so that a voice costs $0, not $15,000.
What it does
Svara is an on-device speech interpreter that turns unintelligible speech into clear voice:
- Listen — the user speaks into the phone as they can. A quantized Whisper model transcribes locally — audio never leaves the device.
- Understand — a personal phrasebook plus an adaptive fuzzy matcher (phonetic + edit-distance scoring) maps garbled transcripts to likely intents.
- Confirm — the top intents appear as large tap targets; one tap confirms (a caregiver can assist). Every confirmation teaches Svara that user's pronunciation — accuracy compounds with use.
- Speak — the confirmed intent is spoken aloud in a clear synthesized voice.
It runs entirely in the browser — WebGPU where available, WASM everywhere else — works offline after first load, and installs as a PWA on a $100 Android phone.
How we built it
- transformers.js + ONNX Runtime Web — quantized Whisper ASR and MMS-TTS run fully in-browser; we vendored the WASM binaries so the app works with zero CDN dependencies
- Adaptive intent layer — edit distance + phonetic-key similarity scored against both the phrasebook and previously-confirmed pronunciations stored locally ("Svara learns the way you say it")
- Zero backend, zero API keys — a static page, because "offline, $0 marginal cost" is the product, not a constraint
- A simulate button feeds canned slurred transcripts through the real matching pipeline so the interaction loop can be demoed on stage
Challenges we ran into
- Baseline ASR genuinely struggles with severe dysarthria — instead of pretending otherwise, we designed around it: the confirm-and-learn loop is standard AAC clinical practice (partner-assisted input, personalized vocabulary) implemented in software, and our roadmap includes LoRA fine-tuning on dysarthric corpora (UASpeech, TORGO)
- WebGPU init hangs silently on some drivers — we added timeout-based automatic fallback to WASM
- "Offline" had to be real — model files and ONNX Runtime WASM binaries are vendored/cached locally, so the app keeps working in a hospital ward or a village with no signal
Accomplishments that we're proud of
- A working prototype — not slides: mic → on-device ASR → intent match → confirm → speech, all in-browser
- The learning loop works today: confirm a garbled transcript once, and similar pronunciations resolve automatically next time
- An honest architecture for the people AAC fails most: offline-first, multilingual (1000+ languages via MMS-TTS), commodity hardware
What we learned
- The confirmation tap is not a failure state — in AAC practice it is the interface; designing for it as a learning signal turned our biggest weakness into the core feature
- Accessibility constraints (big tap targets, partner mode, offline) aren't edge cases here — they're the product
- On-device inference has quietly become good enough to replace a category of five-figure assistive hardware
What's next for Svara
- Phase 1 — LoRA fine-tune Whisper on dysarthric-speech corpora (UASpeech, TORGO); publish WER improvements
- Phase 2 — voice banking: patients record a short script early in the disease, few-shot neural TTS (F5-TTS/XTTS-class) speaks as them
- Phase 3 — pilots with rehab centers and speech-language pathologists; switch access, eye-tracking, partner mode; localized UI
Built With
- accessibility
- healthcare
- javascript
- machine-learning
- on-device-ai
- onnx-runtime
- pwa
- speech-recognition
- text-to-speech
- wasm
- webgpu
- whisper
Log in or sign up for Devpost to join the conversation.