Inspiration

Imagine knowing exactly what you want to say, but having to repeat it until someone understands. Ordering lunch, answering a phone call, or telling someone you love them can become an exhausting negotiation.

For people with dysarthria, difficulty controlling the muscles used for speech can make their words hard to understand. The scale is significant: 22–58% of people with acute stroke present with dysarthria, according to the American Speech-Language-Hearing Association. Beyond stroke, up to 89% of people with Parkinson’s experience speech changes, which can reduce confidence and participation in conversations.

That inspired VoiceBridge: a communication aid that learns how someone speaks, helps recover their words, and asks for confirmation when uncertain. We want people to participate in conversations with greater independence—and hear their words spoken in a voice that still feels like their own.

What it does

VoiceBridge turns difficult-to-understand speech into clearer spoken communication through three steps:

  1. Recognize: A speech recognition model adapted for dysarthric speech transcribes the recording.
  2. Confirm: When the system is uncertain, it shows possible interpretations and lets the speaker choose.
  3. Speak: A voice synthesis model reads the confirmed words aloud using a voice shaped by a reference recording of the speaker.

Our browser demo shows the original model’s transcript alongside the adapted model’s output, so users can see what our training improved.

The adapted Cohere model is hosted on Baseten, and the demo’s speech recognition inference runs on Baseten. The browser streams audio to that deployed model and receives the transcription before confirmation and voice synthesis.

The current prototype demonstrates this complete pipeline on recorded speech. Reliable everyday conversation is our next goal.

How we built it

We explored two speech recognition models: NVIDIA Parakeet and Cohere Transcribe. We adapted them using TORGO, a research dataset containing recordings from people with dysarthria.

Baseten powered the GPU training runs for both Parakeet and Cohere Transcribe. We used Baseten H100 infrastructure to establish baselines, run fine-tuning experiments, and evaluate the resulting models.

We measured accuracy using word error rate (WER): the number of substituted, missing, or extra words divided by the number of words in the correct transcript. Lower is better.

Our fine-tuning engineering went beyond a standard training run: we redesigned how speakers were sampled, controlled which parts of each model could learn, built stronger evaluation splits, and tested competing adaptation strategies.

Adapting Parakeet

Our first attempt at updating the entire Parakeet model made recognition worse. We switched to small trainable components called adapters, which let the model learn new speech patterns while preserving most of its existing capabilities.

We then addressed a major imbalance: some speakers appeared almost five times as often as others. We balanced training exposure across speakers and accounted for recordings of the same utterance captured by different microphones.

We focused particularly on F01—the dataset’s anonymous identifier for the female speaker whose speech was hardest for our frozen Cohere baseline to recognize among the three female speakers evaluated. We also tracked M01, a male speaker with difficult-to-recognize speech, while checking that improvements did not come at the expense of other speakers.

We compared two approaches: training adapters alone, and training adapters while also allowing the final two acoustic encoder blocks—the parts that interpret speech sounds—to change. Those pretrained blocks learned at one-tenth the adapters’ learning rate to limit disruption. We repeated this comparison with the larger, 1.1-billion-parameter Parakeet model.

Adapting Cohere Transcribe

We expanded Cohere’s adaptation from the text decoder alone into the acoustic encoder, allowing it to learn how dysarthric speech sounds—not just how transcripts are written.

We used LoRA, a method that adds small trainable updates to a mostly frozen model. Our configuration targeted the final six encoder blocks and all eight decoder layers, training approximately 3.42 million parameters—just 0.166% of the model.

We also fixed model-specific issues that could undermine training:

  • Repaired compatibility problems in model loading and text generation.
  • Used the same explicit English-language instructions during training and inference.
  • Corrected training targets so the model learned to predict the next token rather than copy one it had already received.
  • Excluded instructions and padding from the training loss.
  • Compared larger adapters and broader encoder coverage; the smallest tested configuration performed best on validation.

Connecting the product

After training on Baseten, we deployed the adapted Cohere model on Baseten for live inference. A WebSocket connection carries audio from the browser to the model and returns transcription updates.

We connected those updates to a confidence-based verifier and OpenVoice V2 for personalized speech synthesis. The verifier only offers interpretations actually proposed by the recognizer, giving the user control when the words are uncertain.

Challenges we ran into

Our biggest challenge was learning from roughly four hours of unevenly distributed training speech. More recordings did not always mean more independent examples: many were two microphone views of the same utterance.

We also had to distinguish real improvement from memorization. A model can look accurate when evaluated on speech it has already encountered. We separated training and evaluation data, excluded matching utterances and held-out prompts, and reported results for familiar and unfamiliar speakers separately.

More computation did not always help. A Parakeet experiment that searched more possible transcriptions was approximately 25 times slower in batched evaluation without improving F01. Allowing more pretrained layers to change sometimes improved development scores while worsening performance on separate evaluation recordings.

Connecting the models introduced further challenges: cold starts, network delays, speech synthesis time, and browser audio conversion. Even converting audio to floating-point samples and back changed one transcription. We fixed that by streaming the recording’s original audio bytes.

Accomplishments that we're proud of

  • Improved recognition for a speaker excluded from training and model selection. On M02, an anonymous male test speaker with 388 recordings, adapted Cohere reduced WER from 56.68% to 34.81%—a 38.6% relative reduction. The number of completely correct transcriptions increased from 98 to 198.
  • Reduced mistakes on important words. On the same test set, our measured error rate for critical terms, such as negations and numbers, fell from 40.00% to 28.57%.
  • Made substantial progress on a particularly challenging female speaker. For F01, Parakeet 1.1B with adapters and two trainable encoder blocks reduced development WER from 61.36% to 22.73%. Adapter-only training performed better on separate held-out prompts, reducing WER from 77.78% to 36.11%, and reached 26.09% on a four-recording diagnostic set. Our 20–30% target was reached on development and diagnostic evaluations, but not consistently on held-out prompts.
  • Used Baseten to train efficiently. Cohere’s 1,200-step training run took approximately 10.6 minutes on one Baseten H100, with the best checkpoint selected before the end of training.
  • Built a working path from audio to personal speech. Our demo connects Cohere inference hosted on Baseten to transcription, user confirmation, and personalized playback. Seven curated demo recordings transcribed exactly, with initial text arriving in approximately 0.6 seconds in our scripted check. These selected examples demonstrate the pipeline; the full test-set results describe broader accuracy.

What we learned

Data selection and evaluation mattered as much as model size. Improving the average score did not guarantee improvement for the person struggling most, so we needed to measure each speaker separately.

We learned that carefully targeted adaptation can outperform changing more of a model. Larger adapters and longer training were not automatically better, especially with limited data.

We also learned that a falling training loss does not guarantee useful speech recognition. Details such as language instructions and correctly aligned training targets determine what a model actually learns.

Finally, a fast model does not automatically create a fast conversation. Recording, network transfer, repeated decoding, confirmation, and synthesis all contribute to the wait. An assistive tool needs both accurate words and an interaction people can comfortably use.

What's next for VoiceBridge

Our next priority is testing with more speakers and fresh, consented conversational recordings. TORGO contains read prompts, so our current results do not establish performance in spontaneous conversation.

We want to introduce short, guided personalization sessions so VoiceBridge can learn an individual’s speech patterns. We will also compare more adaptation strategies using fresh evaluation data and continue improving recognition of difficult isolated words.

On the product side, we plan to reduce end-to-end latency, improve when the system asks for clarification, and make confirmation easier. We want to work directly with people with dysarthria and speech-language professionals to evaluate clarity, voice identity, and everyday usefulness.

Our goal is a communication aid that learns how someone speaks, helps them express what they mean, and keeps them in control of their voice.

Built With

Share this project:

Updates

Submission history