Inspiration
Picture two people reading the same speech. One rushes through the middle and then freezes before the ending. The other sounds flat the whole way through. If you ask three judges why one of them lost points, you'll hear "it felt rushed" or "it didn't land", and nobody can say which second it happened at. I wanted a tool that could.
What it does
SpeechLens takes a recording and its transcript, and compares it with a reference reading of the same text. For each place where the two differ, it gives the start and end time, the words involved, the value it measured, the value it expected, and a short sentence that uses those numbers. The scores and the sentences come from fixed formulas. There's no language model anywhere in the scoring or detection.
It works in two modes. In paired mode you give it an ideal reading of the same text to compare against. In reference-free mode there's no ideal reading, so it compares against statistics collected from the ideal readings in my development set, and it can only look for fewer kinds of flaws.
The dashboard shows the waveform with clickable flaw regions, the transcript, pitch, loudness and speech rate against the reference, six rubric scores, and a reliability badge on every flagged region.
How I built it
There's no public dataset of good and flawed delivery of the same text, so I made one from LibriSpeech. I picked 16 passages from 16 different speakers, 8 men and 8 women, and split them by speaker into 8 for development and 8 for testing. Each passage is the ideal reading. From each one I made five flawed versions, level 1 to level 5. The flaws I injected are fast and slow pace, long pauses, flattened pitch, volume drop-off, filler sounds and repeated words, at severities between 0.2 and 0.9. I also made four control versions with nothing wrong: an identity pass through the same processing, an MP3 round trip, a 6 dB gain change and some added noise. They tell me how often the detector fires when it shouldn't. That's 160 recordings in total.
Since I put each flaw in myself, I know exactly where it is. The catch is that stretching audio or adding a pause moves every word after it, so each label is saved twice, once for the original timeline and once for the modified audio.
The analysis starts with forced alignment using torchaudio's MMS model, so every word has a start and end time. Then I compute pitch, loudness, speech rate and pauses every 10 ms. Pitch is in semitones relative to the speaker's own median and loudness is in dB relative to their median, so a deep voice and a high voice get compared fairly. Every flaw type has its own detector and its own threshold. I tuned those on the development speakers only, froze them with a checksum, and ran the test speakers once.
It's written in Python with PyTorch, torchaudio, Praat (through Parselmouth) and librosa. The backend is FastAPI, the dashboard is plain JavaScript with Plotly, and tests run on GitHub Actions.
Challenges I ran into
My first results looked great, and they were wrong. The detector was reading word timings from the dataset's label files, and those files already included the fillers and repeated words I'd injected. It was reading the answers. When I switched to forced alignment on the audio alone, the F1 score dropped from 0.53 to about 0.05, with dozens of false alarms a minute.
After that I spent most of my time debugging the measurements. One of the pace measurements cancelled out the very stretch it was supposed to catch. The pitch measure ignored moderate flattening. The filler score never ran in paired mode. A single threshold for everything didn't work either, so each detector got its own, and each mode got its own false-alarm budget.
There were smaller problems too. At 16-bit, the quiet room noise I used to fill inserted pauses rounded down to pure silence, which any detector could spot trivially, so I stored the dataset as 24-bit. And Praat's resynthesis gives slightly different output on each run, so I cache it to keep builds repeatable on one machine.
Accomplishments that I'm proud of
I'm proud of the dataset, with its two-timeline labels, control versions, checksums and datasheet, and of sticking to the order of tune, freeze, then test once. I'm also glad I kept the weak results in. On the test speakers, paired mode reaches an F1 of 0.383 (95% interval 0.311 to 0.453), compared with 0.464 on development. Long pauses do best at 0.79 and fillers get 0.67. The total score drops as the flaws get worse, with a Spearman correlation of -0.89 on the test speakers.
What I learned
A shortcut in how you load your data can hide a leak for a long time, and the unflawed control recordings turned out to be as useful as the flawed ones. I also learned that a lower number you can defend is worth more than a higher one you can't.
What's next for SpeechLens
There's a lot that doesn't work yet. Slow pace, monotone and volume drop-off have low precision. A 6 dB gain change causes false regions in paired mode, 3.09 per minute. Region edges are about 256 ms off on average. On the test speakers, false alarms came to 1.05 per minute, just over my budget of 1. Repeated-word detection is switched off because it needs speech recognition. And since all the flawed recordings are synthetic, I haven't checked any of this against real people making real mistakes.
Next I'd want a comparison that isn't thrown by loudness changes, a speech recognizer for repeated words, and a small set of real recordings that people have annotated by hand.
The full technical report (4 pages, PDF) is here: REPORT.pdf
Log in or sign up for Devpost to join the conversation.