Inspiration
Delivery feedback is usually vague: "work on your pacing", "more energy". Two judges can hear the same rushed phrase and disagree on where it started and how bad it was. And no public dataset pairs a good delivery with bad deliveries of the same words, so there's no way to measure whether a speech evaluator is actually right. We wanted feedback that points at an exact moment and says why.
What it does
Pick a reference excerpt from a great public-domain speech (JFK, Reagan). Then record yourself reading it in the browser, upload a file, or try one of the dataset samples. Podium returns:
- a deterministic 0-10 rubric for pacing, pauses, pitch, volume and fluency;
- flaw regions with exact start and end times, shaded on time-series overlays of your pitch, loudness and pace against the reference, time-warped onto your timeline word by word;
- a causal explanation for every region, built from the measured numbers ("14.3 syllables/s here vs 7.7 in the reference, +85%"), plus a concrete tip. Click a flaw to hear it.
Your overall style is reported once ("overall 12% slower than the reference") and kept apart from local flaws, so a different voice isn't flagged on every sentence.
How we built it
Dataset (207 recordings, labelled to the millisecond). Three public-domain speeches: JFK's 1961 inaugural and 1962 Moon speeches, and Reagan's 1986 Challenger address. Whisper large-v3-turbo transcribed applause-free windows, and the MMS wav2vec2 forced aligner timed every word. We then built the "bad mirror": 8 flaws (rushed, dragged, monotone, mumbled, volume spike, awkward pause, missing breath, stutter) at 3 severities, injected into the same audio. We used phase-vocoder time-scaling, WORLD-vocoder pitch flattening, filtering, room-tone silence and onset repetition. Edits keep segment lengths exact, so every label is sample-accurate. Each baseline gets a gradient from L0 (untouched) to L4 (egregious), 24 single-flaw stress tests, and readings by open Piper TTS voices to test that it works across speakers.
Analysis. Both recordings are force-aligned to the same transcript, so word i matches word i. Praat (Parselmouth) and librosa extract speaker-normalised features: pitch in semitones relative to the speaker's median, loudness relative to their own level, articulation rate, pauses, high-frequency consonant energy, and voiced sound inside gaps. That last one separates a stutter's restart from a silent pause. Each word gets two robust z-scores: z_ref (departure from the reference beyond the speaker's overall style) and z_self (standing out within their own delivery). The flaw score is min(z_ref, z_self). Flagged words merge into time regions, and each region is explained from its own numbers.
Evaluation. Thresholds were tuned only on the JFK speeches, then frozen. Reagan (a different speaker and speech) and an unseen voice are the held-out test. On it:
- mean overlap (IoU) with the true flaw: 0.81; overall F1 0.51;
- pauses, mumbling and stutters located within 0.00-0.06 s;
- zero false alarms on untouched recordings of the reference speaker;
- the rubric falls monotonically from L0 to L4 on every test excerpt (Spearman ρ = -1.00).
The app is a FastAPI backend with a single-page Plotly dashboard and in-browser recording. Everything runs locally.
Challenges we ran into
- Whisper's long-form mode hallucinated 20 seconds of "Bye-bye" over applause, so we switched to transcribing clean windows one at a time.
- Half-precision Whisper silently returned empty text on our GTX 1660 Ti; only float32 worked.
- The first version flagged a different speaker on nearly every sentence (41 false alarms per minute). The fix was conceptual: separate style (consistent differences) from flaws (local departures), and require both kinds of evidence. That cut false alarms to 3 per minute.
- A stutter's restart often lands outside the aligned word, so it looked like a pause. Measuring voiced sound inside gaps fixed it.
Accomplishments that we're proud of
- A contrastive dataset where every flaw label is exact, not a guess, published with code to rebuild it.
- A real held-out test (new speaker, new speech, unseen voice), not just numbers on the data we tuned on.
- Explanations that are numbers, not adjectives: every flaw says what was measured and how far it was from the reference.
- Honest reporting, including what doesn't work yet.
What we learned
Most of the difficulty isn't detecting a deviation. It's deciding which deviations count. A reference delivery is one style, not the truth, so a good evaluator has to separate "different from JFK" from "worse than yourself here". We also learned to trust evaluation over intuition: several ideas that looked right in a single example made the held-out numbers worse.
What's next for Podium
- Pacing: it's the weakest flaw type (F1 0.29-0.35). Slowing a phrase also stretches its pauses, so drags are sometimes reported as awkward pauses. A syllable-nucleus detector should help.
- Different speakers: F1 is 0.21 vs 0.61 for the same speaker. Next is learning noise models from paired readings instead of fixed floors.
- Human data: human-recorded flawed readings with human labels, collected through the dashboard's record button.
- References: multiple references per text, so "ideal" becomes a range rather than one person's style.
Log in or sign up for Devpost to join the conversation.