Inspiration
As soon as we saw the NSA Audio Authentication challenge, both of us knew that this was the problem we wanted to tackle at HackGT 13. We are both very interested in machine learning and cybersecurity, and this was such a cool intersection of both of these things. Plus, the problem is huge, highly relevant today, and critical to national security.
What it does
This project is our best attempt to generate an accurate probability that any inputted audio file (English, at least 2 seconds) features a synthetic or AI-cloned voice rather than a real human one.
How we built it
First off, we did a lot of research on voice cloning and synthetic audio detection. Then, we decided on which indicators we felt would be the most helpful/important to detect voice cloning, came up with rough ideas of how we wanted to combine and weigh them in our final scores, researched the machine learning models, open-source tools, etc. that would be most useful to actually evaluate the indicators or features, and then started coding from there!
Challenges we ran into
I think one of our biggest challenges was that we wanted to use a lot of novel technologies in this project, such as quantum machine learning. However, we had to then reprioritize and focus on what technologies and models were actually improving our accuracy.
Accomplishments that we're proud of
- We are really proud that we had some conditional logic in our algorithms rather than pure brute-force or the same approach for every instance. We use an unsupervised K-means condition router to inspect signal-level recording quality indicators and dynamically adjust model-blending weights, ensuring noisy or truncated clips are evaluated under optimized, cluster-specific rules.
- We also love that we were able to avoid black-boxing by returning detailed JSON for each clip with a score breakdown and details on exactly which variables led the decision tree to push it towards synthetic or real.
What we learned
We learned so, so much about voice cloning in general and how it works. We were able to strengthen our machine learning and programming skills through this project and also enter an AI space that we were very unfamiliar with before.
What's next for Identifying Voice Cloning
In the future, we definitely want to add in more indicators or ways to identify voice cloning, such as deep neural anti-spoofing networks, phase break splice detectors, acoustic environment (ENF mains-hum) matchers, and speaker-embedding drift tracking. We also look forward to looking at other teams' implementations and learning/getting ideas from those.

Log in or sign up for Devpost to join the conversation.