Synthesis

Inspiration

In Public Forum Debate, judging is bound to grow inconsistent; the same round can get wildly different scores based on who’s judging. To solve this, we created Synthesis, a debate judging tool that gives judges a consistent, explainable second opinion without overruling human decision-making. Whether a judge missed a contention or couldn’t understand a refute, Synthesis is here to give not only students a fair, valid opinion, but judges the right stats to use.

What it does

Both Judges, for aid in scoring, and Students in need of critique can upload or record their debate speech, and then our app returns a full breakdown:

Delivery analysis: pitch, jitter, shimmer, energy, vocal clarity, and projection, extracted directly from the audio waveform

Argument extraction: every contention is broken into Assertion / Reasoning / Evidence / Impact (AREI), with a 1-5 validity score on whether the evidence actually supports the reasoning, and whether the reasoning actually supports the assertion

Evidence detection: flags whether claims are backed by evidence or left as bare assertion

Refutation analysis: detects rebuttal moments and rates how well each one actually dismantles the point it's responding to, rather than just acknowledging it exists

Format compliance: checks contention counts and speech length against Public Forum debate rules

Video analysis: eye contact and posture stability from facial and body landmark tracking

Speaker points: a combined 0-82 score weighing all of the above

How we built it

We deliberately avoided wiring together a pile of paid emotion/sentiment APIs and calling it a day. Instead:

Audio features (pitch, jitter, shimmer, energy) come from parselmouth, a Python wrapper around Praat; the same tool used in actual speech-science research- not a black-box emotion API

Speech-to-text runs locally via faster-whisper, no external API calls Emotion classification is our own model: trained on the RAVDESS dataset, using hand-extracted acoustic features (pitch, MFCCs, energy, zero-crossing rate), with emotion categories collapsed from

RAVDESS's 8 acting emotions down to 3 debate-relevant ones (assertive/composed / hesitant) — and validated with an actor-held-out train/test split so the model can't just learn to recognize whose voice it's hearing

Contention detection uses sentence embeddings (all-MiniLM-L6-v2) and a rolling-window cosine-similarity comparison to find real topic shifts in a transcript- not keyword matching

Argument extraction and validity scoring is the one place we use an LLM (Claude), and only for a narrow, structured task: extract four labeled fields from one contention, and separately rate two specific logical relationships- never "just grade the whole speech". We do this to save time, usage, and to prevent hallucinations

Video analysis uses MediaPipe's face and pose landmark models to compute eye-contact ratio and posture stability from real tracked coordinates

Challenges & What We Learnt

Getting jitter/shimmer/clarity numbers that were actually meaningful, not just plausible-looking

Tuning contention-boundary detection so it reflects real argument structure instead of over- or under-segmenting a transcript

Making sure evidence and refutation detection generalize beyond exact keyword matches, using semantic similarity as a second layer

Ran into fragmentation and multiple contention and refutation counts; so we increased the threshold amount, and also the minimum word intake.

Making the video-landmark overlay fast enough to not blow up processing time, and encoding it in a format browsers will actually play.

Landmarks weren’t accurate to what was actually required, making prediction confidence weak, so we adjusted to have better focal scope on the eye tracking.

What's next

Collecting real debate audio to retrain the emotion classifier instead of relying solely on RAVDESS

POI/heckle tracking (currently scoped for manual tagging, not yet built)

Empirically validating the speaker-points formula weights against real judge scores

Built With

Share this project:

Updates

Submission history