Inspiration

A pause is not always the end of a thought.

For someone who stutters, a pause can mean they are still working through a word. A voice assistant that treats that pause as the end of a turn can start answering before the person has finished speaking.

We wanted to explore how voice AI could give people more room to finish without simply making everyone wait longer.

Our first exploratory evaluation gave us a reason to investigate: stock Smart Turn v3.2 INT8 predicted completion at its first check in 116 of 200 automatically detected gaps in stutter-labelled recordings. That measured model decisions, not audible interruptions—but it showed us where to look.

StutterTurn grew from a model experiment into a working voice-agent system. We wanted to hear what happens after a premature decision, compare our adapted model against stock configurations, and make those differences visible.

What it does

StutterTurn lets developers replay real speech through different turn-taking configurations and inspect when the agent decides to answer, waits, cancels a response, or starts speaking.

Our hosted console supports stock models, a longer-wait configuration, and our exploratory adapted model, X1. Visitors can select a recording, configuration, and voice, then hear the source recording and generated response directly in their browser.

The entire backend runs on Vultr. It performs turn-model inference, sends captured audio to Gemini for a reply, generates speech with ElevenLabs, and streams audio and events to the browser.

Our judge page presents two sets of Stock FP32 versus X1 comparisons:

  • A longer, 10.4-second recording that highlights an early interruption.
  • Five additional recordings with matched timelines for both configurations.

In the selected longer replay, stock decided to respond during an early pause. The speaker resumed, and stock reported 0.050 seconds of reply-audio overlap. Our X1 configuration held that pause and reported no overlap.

Both configurations used the same 0.2-second silence gate, but different model weights and decision thresholds. This is a concrete configuration-level result on one recording, rather than proof of a general reduction in interruptions.

Tiger Data powers the saved comparison evidence. Judges can connect the displayed timelines to the underlying run logs and SQL queries, then launch a fresh audible replay.

The current demo uses recorded speech. Microphone and telephone input are future extensions.

How we built it

We built three connected pieces: a model-adaptation workflow, a cloud-hosted replay agent, and an evidence interface.

Adapting the turn-taking model

We started with Pipecat’s Smart Turn v3.2 and real podcast speech referenced by SEP-28k.

Using the source-pinned community haan converter, we recovered the released FP32 weights into a trainable PyTorch model. We checked numerical agreement with the original ONNX model, verified training updates, and tested export back to ONNX.

We compared stock INT8, stock FP32, longer silence waits, threshold adjustments, and adapted candidates. We also checked for regressions on other speech, because reducing premature decisions on one group of recordings is not enough if the model becomes worse elsewhere.

Our current X1 candidate remains exploratory.

Google Gemini API: generating the response

When the turn-taking policy decides to answer, the backend sends the captured audio to Gemini.

A dedicated system instruction asks for a brief, useful response, requests clarification when needed, and avoids unsolicited commentary about how the person speaks.

Gemini is part of the actual hosted interaction: its generated text becomes the voice agent’s spoken reply.

ElevenLabs: making the interaction audible

We send Gemini’s reply to ElevenLabs streaming text-to-speech.

The backend forwards the generated PCM audio over an authenticated WebSocket, and the browser plays it using the Web Audio API. Visitors can hear the source recording and the agent’s response, adjust volume, stop playback, or reset the run.

Cancellation is part of the interaction. When the speaker resumes, the system can cancel a pending or active reply. This lets us investigate whether an early model decision actually becomes spoken overlap.

Tiger Data: explaining what happened

We store run metadata in PostgreSQL and timestamped events in a TimescaleDB hypertable.

SQL views reconstruct completion decisions, resumed speech, cancellations, and response timing. Our evidence database contains 37 runs and 631 events, including earlier failed attempts and superseded runs.

The judge page uses saved query exports for the selected comparisons. We verify those exports against the original logs and keep each comparison tied to specific run identities.

Tiger Data gives us a way to answer questions such as: Did the model decide too early? Did the speaker resume? Was the reply cancelled before audio? Did generated speech overlap the resumed turn?

Vultr: hosting the working system

Vultr hosts our HTTPS website and Python backend, including ONNX inference, Gemini requests, ElevenLabs requests, and the HTTP/WebSocket API.

Caddy serves the site and proxies the backend, while systemd manages the service. Provider keys stay server-side, demo controls require authentication, and the instance admits one replay at a time.

The frontend uses HTML, CSS, and JavaScript. The backend uses Python, Silero VAD, and ONNX Runtime.

Challenges we ran into

One of our biggest challenges was separating a premature decision from an actual interruption.

A model can decide to answer too early, but the speaker may resume before response generation finishes. If the system cancels in time, no agent speech is played. We had to track decisions, resumption, cancellation, and audio separately to understand the outcome.

Another challenge was competing with a simple solution: waiting longer. A fine-tuned model has to justify itself against a longer silence timeout, so we included that baseline in our experiments.

Model adaptation also produced tradeoffs. Some candidates improved development results while regressing on other unfinished speech. Our earlier v1–v4 attempts did not pass their selection checks. Those results helped us investigate input duration, recording differences, and automatically generated labels.

Moving the demo from a laptop to Vultr introduced another problem: a server can generate audio without delivering it to a visitor’s speakers. We built browser audio transport and checked that both source speech and generated replies rendered in Chromium.

We also encountered empty Gemini replies and real request timeouts. We corrected the error classification, added one bounded retry for empty responses, and increased the reply deadline from 10 to 20 seconds. These changes improve recovery, although live API availability and latency remain external dependencies.

Accomplishments that we're proud of

We built a complete hosted interaction, from turn-model inference to a generated spoken reply in the browser.

All 12 runs selected for our current Stock FP32/X1 comparisons completed with generated speech rendered in Chromium.

Our longer replay demonstrates the behavior we set out to investigate: stock answered during an early pause and reported a brief overlap, while X1 held that pause and reported none. Both configurations later cancelled another premature reply before audio, showing why cancellation also matters.

Our model appears alongside stock across all five additional comparison recordings. The timelines distinguish model-triggered completion from the three-second silence fallback, so waiting behavior remains visible.

We established a verified path from released ONNX weights to training and back. Our reconstruction and export checks matched reference probabilities within approximately 1.2 × 10⁻⁷ on synthetic checks.

We also connected model identities, configurations, run logs, database queries, and displayed results. Forty-seven backend tests passed locally and on Vultr, alongside browser checks for audio, cancellation, authentication, and playback state.

X1 was selected after reviewing earlier development results and remains an exploratory candidate. These replays demonstrate specific observed behavior; they do not establish general model superiority or measured acoustic interruption rates.

What we learned

Turn-taking is a property of the whole system.

The model makes a decision, Gemini generates a response, ElevenLabs produces speech, and the playback system starts or cancels that speech. Evaluating only the classifier misses much of what the speaker experiences.

We also learned that patience must be evaluated alongside responsiveness. A system that avoids interrupting by waiting too long can create a different frustrating experience.

Finally, traceable evidence made our project stronger. Keeping failed experiments and provider failures helped us understand the system and explain exactly what our working demonstrations show.

We build on Smart Turn, Whisper, Silero VAD, haan, and SEP-28k, and credit their creators and the original podcast speakers. Our examples use real recordings; nobody imitates stuttering. Claude Code and Codex assisted with implementation, documentation, and review.

What's next for StutterTurn

We want to move from promising replay behavior to stronger evidence of useful turn-taking.

Our next steps are to review speech boundaries with human annotators, freeze a confirmation set, and evaluate X1 against stock and longer-wait baselines. We will measure premature responses, completed-turn response delay, cancellations, and nonresponses together.

We also plan to automate ingestion of new Vultr runs into Tiger Data, expand browser and listening checks, and test longer recordings with more speakers.

From there, we want to add microphone and telephone input and work with people who stutter to understand their interaction preferences.

Our goal is simple: voice AI that gives people room to finish.

Built With

Share this project:

Updates

Submission history