From there we shifted the primary signal from what the voice sounds like to how the conversation happens: turn-taking dynamics (response latency, barge-ins, speech ratio) turned out to be both cheaper to compute and more robust to unseen synthetic voices, since conversational timing generalizes across TTS engines better than acoustic fingerprints do. Combining 11 turn-taking + acoustic features in a logistic regression model got us to 95.8% validation accuracy (AUC 0.967), with sub-3-second latency per call, validated on a speaker-disjoint train/val split so no voice leaks between sets.
Log in or sign up for Devpost to join the conversation.