Inspiration
About 1 in 14 U.S. children ages 3–17 has a voice, speech, or language disorder, yet fewer than 60% receive intervention services.
We built Vocally, a multimodal AI speech therapist that combines conversational AI with real-time speech analysis to help users practice, understand, and improve their speech between traditional therapy sessions.
Unlike a normal voice assistant that only understands what you say, Vocally also analyzes how you say it through a custom speech-classification model.
What it does
Vocally is a multimodal AI speech therapist that combines real-time conversation, speech analysis, and guided practice into one experience.
Users can:
- Conversational AI — Practice speaking through natural, supportive conversations
- Two-Headed Speech Model — Analyze whether a speech event occurred and what type it was
- Conversation — Speak freely while Vocally listens and adapts in real time
- Speech Exercises — Practice structured, evidence-based speech techniques
- Endless — Build confidence through a speaking game that rewards meaningful communication
- Provider Reporting — Export patient progress and speech data for healthcare providers upon request
Our goal is simple: amplify therapy, not replace it.
How we built it
Conversational AI
Vocally streams 16 kHz audio through two parallel paths: one understands what you said, while the other analyzes how you said it. Gemini handles conversation, and ElevenLabs streams the response back at 24 kHz.
Two-Headed Model + Pipeline
We replaced a ~350M-parameter open-source model with our own 1.3M-parameter CNN — about 269× smaller.
$$ \frac{350M}{1.3M} \approx 269\times $$
Inference dropped from roughly 300–350 ms to 5–7 ms, while our final classifier reached 0.823 macro ROC-AUC across 6 speech patterns.
Conversation
Each speaking turn combines:
$$ \text{Meaning} + \text{Speech Patterns} $$
The conversational model tracks context while our classifier detects blocks, prolongations, repetitions, interjections, and fluent speech in real time.
Speech Exercises
We researched a framework containing 27 speech techniques and implemented exercises including Syllable-Timed Speech, Cancellation/Pull-Out, and Preparatory Set.
Research on Syllable-Timed Speech has reported reductions in stuttering of up to 96% in the studied setting.
$$ \text{Improvement} = \frac{\text{Baseline} - \text{After Practice}} {\text{Baseline}} \times 100 $$
Endless
Endless rewards meaningful communication instead of fewer stutters.
$$ \text{Score} = \left( \frac{C}{T} \times 100 \right) (1 + R) $$
Where:
- C = content words
- T = speaking time
- R = sentence richness
Repetitive filler can score 0, while about 90 seconds of meaningful speech can score 300+.
Provider Reporting
Vocally tracks speech events, exercise performance, and progress over time into a structured practice report — frequency scored against the SSI-4 framework clinicians already use, so a provider can read it in language they recognize, not a format we invented.
It's exportable on request and meant to sit alongside a clinician's own assessment, not replace it.
$$ \text{Patient Data} \rightarrow \text{Structured SSI4 Report} \rightarrow \text{Healthcare Provider} $$
This lets Vocally bridge the gap between at-home practice and professional care without trying to replace the clinician.
Challenges we ran into
The biggest challenge was balancing accuracy and latency. Our 350M-parameter model performed better on paper, but it made the conversation feel noticeably slower. Even a few hundred milliseconds of silence can make a voice assistant feel unnatural.
Building our own CNN taught us that the best model is not always the biggest one. It is the one that actually fits the product.
The second challenge was more human than technical. Speech feedback can be sensitive, and we did not want Vocally to make someone feel like they had failed just because it detected a stutter.
Our classifier reached around 0.82 AUC, which is strong, but not perfect. Because of that, we chose not to show raw speech labels during a live conversation. That decision also inspired Endless, where we reward how much someone communicates instead of how fluently they speak.
Finally, building the real-time pipeline was difficult. We had audio coming in, speech classification, transcription, Gemini, and streamed ElevenLabs audio all happening at once.
These pieces had to run concurrently so our 5 to 7 ms classifier latency actually mattered.
Accomplishments that we're proud of
- Replaced a 350M-parameter open-source model with our own 1.3M-parameter CNN, making it 269x smaller and 47x faster
- Reached 0.823 macro ROC-AUC across six speech patterns
- Built a two-stage model that first detects a speech event, then determines its type
- Built a real-time pipeline combining audio analysis, Gemini, and streamed ElevenLabs speech
- Grounded our exercises in real speech therapy techniques
- Created Endless, a scoring system that rewards expression instead of penalizing stutters
- Deployed the entire system end to end, rather than stopping at a local prototype
What we learned
We learned that building health-focused AI involves more than getting the model to work. Some of our most important decisions were about when not to use the model's output.
We could show every classification in real time, but that does not mean we should. Thinking about how a wrong prediction could affect the person using Vocally changed how we designed the entire experience.
We also learned how important latency and concurrency are in voice AI. A model can be accurate and still be unusable if it makes the conversation feel slow.
What's next for Vocally
- Image processing to understand facial expressions and visual signs of discomfort
- Provider-ready reports using healthcare standards such as HL7 FHIR
- Personalized models that adapt to an individual speaker
- More evidence-based speech exercises
- Multilingual support
- Privacy-first architecture with local inference and locally hosted patient data
Built With
- elevenlabs
- express.js
- gemini
- javascript
- node.js
- onnx
- python
- pytorch
- react
- tigerdata
- typescript
- vite
- websockets
Log in or sign up for Devpost to join the conversation.