Inspiration

The NSA's HEARSAY challenge put a real problem in front of us: as AI voice cloning and text-to-speech get frighteningly good, how do you tell a genuine human voice from a synthetic one? Voice-based fraud, disinformation, and impersonation are no longer hypothetical — they're here. We wanted to build a detector that could catch fakes reliably enough to actually be trusted.

What it does

HEARSAY classifies audio files as genuine human speech or AI-generated/synthetic speech. Given a set of recordings, it assigns each one a confidence score and separates the real voices from the fabricated ones, the kind of forensic screening you'd want before trusting a voice on a call, a recording, or a piece of evidence.

How we built it

We ran everything on GPU-accelerated infrastructure (NVIDIA H100s on Georgia Tech's PACE cluster). Our first approach extracted features from all layers of a WavLM speech transformer, combined them with classic acoustic features (LFCCs), and trained gradient-boosted classifiers in an ensemble. When that overfit, we pivoted to a fine-tuned WavLM-large model with AASIST graph-attention pooling, trained on millions of speech samples spanning many different synthesis methods. We also curated a diverse real-speech dataset (LJ Speech, LibriSpeech, and thousands of People's Speech samples) to make sure the model generalized across many speakers rather than memorizing one voice.

Challenges we ran into

Our biggest lesson came from a trap: our first model hit 99%+ validation accuracy but performed terribly on the real test set. It had memorized the fingerprints of specific speech generators instead of learning what actually makes a voice synthetic. We also discovered that over half of our "real" training audio came from a single speaker, which skewed the model badly. Diagnosing why "perfect" validation numbers meant nothing on unseen data, and fixing it was the hardest and most valuable part.

Accomplishments that we're proud of

We took our detector from worst place (minDCF 0.913, 34.6% error rate) to a clean, confident classifier with sharp separation between real and fake. We built a full GPU pipeline from scratch, solved the overfitting and data-diversity problems, and produced a model that makes decisive, well-calibrated predictions on audio it has never seen.

What we learned

  • High validation accuracy is a lie if your data leaks. Generalization is everything.
  • Data diversity beats data volume. One dominant speaker can quietly wreck a model.
  • The right pre-trained model beats a hand-built pipeline — knowing when to stop engineering features and leverage a strong foundation model was a key judgment call.
  • How to use limited compute and time wisely under hackathon pressure. ## What's next for hearsay_hgt13
  • Multi-window analysis to scan entire recordings, not just the opening seconds, so fakes can't hide later in a clip.
  • Ensembling multiple detectors for even more robust scoring.
  • Real-time detection so HEARSAY could flag synthetic speech live, during a call or stream.

Built With

Share this project:

Updates

Submission history