Motivation
Track C asks for evidence behind speech feedback rather than opaque grades. CadenceLab makes timing and acoustic differences inspectable, reproducible and explicitly limited.
What it does
A local dashboard compares pre-aligned WAV recordings and a shared transcript. It highlights additional silence, relative energy deviations and clipping with timestamps and mathematical explanations. Genuine Wav2Vec2 CTC forced alignment produces word spans; optional Kaldi-compatible MFCCs add spectral evidence. Audio is not persisted by the server.
How it was built
Python, NumPy, PyTorch and TorchAudio implement 50 ms acoustic frames, RMS, FFT centroid, zero crossings, clipping, MFCCs and model alignment. JavaScript renders time-series overlays and word evidence. A published engineering rubric uses 45% silence, 40% energy and 15% clipping penalties. These weights are not human-calibrated rhetorical judgments; qualityScore remains null.
Dataset and stress testing
Public-domain JFK Rice University speech supplies three disjoint excerpts. The distributed dataset contains 24 recordings, including 21 controlled variants: graded silence and attenuation plus a global-gain control. Provenance, SHA-256 hashes, intervention labels, transcripts, actual CTC alignment and reproducible evaluation are included in the source/dataset archive.
At a frozen 6 dB threshold and one-to-one same-kind temporal IoU of at least 0.5, development precision/recall is 1.0/0.8333; validation and held-out excerpts each achieve 0.8333/0.8333. All excerpts use one speaker. These are controlled signal tests, not independent human speech-quality validation.
Verification and engineering findings
55 local tests passed. Actual browser uploads rendered 48 word rows, MFCC evidence and a silence interval at 6.35-7.35 seconds; downloaded JSON export was checked. Global gain should not be mistaken for a delivery flaw. Moderate attenuation can be missed, while fragmented detections can create false positives; both outcomes are retained in the report.
Limitations
Recordings must already share a timeline. Forced alignment does not time-warp independently timed speech. Transcripts and word boundaries remain provisional. Recording noise can cause false positives. No claim of semantic evaluation, independent-speaker generalization or validated rhetorical quality.
Deliverables and disclosure
Source, tests, paired dataset and two-page report. Extract CadenceLab-source-dataset.zip for the complete dataset and runnable source. Model weights are excluded; official local setup is documented.
4:50 walkthrough uses actual application captures with synthetic Windows narration, not continuous screen recording. Implementation, testing support and documentation were AI-assisted by Codex. AI tool use is disclosed in the README and Built With. No competition acceptance or earnings claim.
Log in or sign up for Devpost to join the conversation.