Inspiration

A modern voice-cloning system needs about three seconds of audio to convincingly impersonate someone. The scam that follows is brutally simple: a grandparent picks up, hears their grandchild's voice saying they're in trouble and need money now, and wires it before anyone thinks to ask "is that real?" The people targeted are usually the least equipped to spot it, and the advice we give them — "listen carefully for something off" — is advice that stopped working.

We also didn't want to build another deepfake detector. Detection is an arms race, generators improve every month, and a family shouldn't have to depend on winning that race.

What it does

Verity protects a family in three layers, in order of how much we trust them.

Verify is the core, and it doesn't use AI at all: a shared safe-word, two or three private challenge questions, and a trusted callback number. A perfect clone can copy how someone sounds but not what only your family knows. This keeps working when the internet is down, the model fails to load, or the detector becomes obsolete.

Detect is the extra signal. Put the call on speaker and Verity analyses it live, showing a scrolling spectrogram and a verdict strip so you can see when something sounded synthetic. It's always framed as evidence never proof — the UI never says "this voice is fake", only "synthetic characteristics detected".

Protect is the action: one tap to verify the caller, one to call back on a trusted number, and one to alert a family member by real SMS.

The same forensic engine answers NSA's HEARSAY challenge. Upload any audio file and it returns a 0–1 synthetic probability plus a plain-language account of which techniques fired and which had no effect.

How we built it

A FastAPI backend and a Next.js frontend, designed for adult children and their parents rather than security engineers — large type, high-contrast and text-size controls, and a simplified "call mode" with one big button.

The forensic engine stacks three frozen self-supervised speech models (WavLM-base, XLS-R 300M, WavLM-large) with classical signal forensics — spectral statistics, prosody, ENF mains hum, compression traces, splice and phase-break detection — and container/metadata analysis. Everything is fused by a regularised linear model selected using NSA's own metric: ASVspoof5 minDCF with π_spoof = 0.3 and C_fa = 4, which we reimplemented and verified against the official scorer to six decimal places.

The system also decides what to run rather than brute-forcing everything: it picks a decoder by container, skips silent windows, samples long files, and runs a cheap model first, escalating to the expensive one only when the cheap one is uncertain.

Challenges we ran into

The training data leaked format, not synthesis. The provided real and synthetic clips differed in sample rate, length, loudness and encoder tags — every real file carried an FFmpeg tag and no fake did. A model could score near-perfectly by learning the file format, then collapse on NSA's normalised test set. We measured the test distribution and rebuilt the entire training set to match it: resampled, cropped, low-passed and peak-normalised, then augmented both classes identically with noise, telephone bandwidth, codecs and reverb. Several forensic cues that had looked strong dropped to chance afterwards — which told us they were never detecting synthesis at all.

Off-the-shelf deepfake detectors barely worked. We integrated five published models and they scored near chance (minDCF 0.946) on modern diffusion and zero-shot generators. They were also 80% of our compute. We dropped them from the model and kept them in the UI as analyst evidence.

Everything ran on a CPU-only laptop. No GPU, so fine-tuning was off the table; we leaned on frozen representations and spent the compute budget on feature extraction and honest cross-validation instead.

Twilio's free tier fought us. Trial accounts can't send custom SMS text — only predefined templates. We built the integration properly, made the template a one-line config switch, and added a "text them yourself" fallback that needs no account at all.

Accomplishments that we're proud of

Cutting minDCF from 0.253 to 0.159 across four model generations, with EER down to 6.0% — measured the hard way, on generators the model had never seen during training.

Validating honestly. We report two protocols: leave-generator-out (pessimistic, drives model selection) and seen-generator (minDCF 0.028, closer to what we believe the test set looks like). We didn't pick the flattering number and call it the result.

A Docker image that we verified with networking completely disabled, reproducing our host predictions to within 0.000007 on all 1,671 test files.

And an app that still protects a family when every AI component fails.

What we learned

Call screening without an app. Calls forward to a Verity number, Twilio streams the audio to our backend, and we score it live and whisper a warning that only the person being called can hear. That works on any phone, including a grandparent's flip phone — which matters, because the people most at risk are the least likely to install anything.

Attack types we don't cover yet. We've started adding replay attacks (a clone played through a speaker into a microphone) and scene manipulation, both of which NSA lists and neither of which appeared in our source data.

Fine-tuning the front-end on a GPU, which is where the remaining accuracy almost certainly is.

What's next for Verity

Call screening that works on any phone. Today someone has to put the call on speaker next to a laptop. Next, calls forward to a Verity number: Twilio streams the audio to our backend, we score it every couple of seconds, and if risk climbs we whisper a warning only the person being called can hear — "this voice may be synthetic, ask for your safe-word." The scammer hears nothing, and it needs no app and no smartphone, which matters because the people most at risk are the least likely to install anything.

On the model side, our weakest categories are the ones our data barely covers: voice conversion, partial splices, and replay attacks — a clone played through a speaker into a microphone, which is exactly what happens when someone holds up a phone to demo it. We've started generating those, and we want real bona fide speech from cars, kitchens and cheap handsets rather than audiobook narrators, since false alarms are what would make a family stop trusting the product. Fine-tuning the speech front-end on a GPU is where the remaining accuracy almost certainly is. But the honest truth is that the safe-word does more work than the detector and needs no technology at all — so the best thing we could build next is whatever helps a family agree on one and practise it, and never need to open the app.

Built With

Share this project:

Updates

Submission history