Inspiration

This project was inspired by issues our parents face as immigrants from China. Oftentimes, when calls need verification of their identities and require certain phrases to do so, our parents struggle to get these systems to understand their words because of the heavy accent they have when speaking in English. While a person face-to-face would be able to understand their words, chatbots on calls, especially with lower audio quality, struggle to do so.

As we researched more into this problem, we realized how impactful a solution would be not just to our parents, but also to the rest of the world. Around 3 million people live with dysarthria, a disorder that affects speech following a stroke, or as a result of Parkinson’s disease, ALS, cerebral palsy, or multiple sclerosis. This condition results in poor articulation that causes trouble for these standardized systems.

What it does

Timbre is a payment approval system that is able to identify the person speaking by the unique nature of their voice, not by what the machine understood from them. This system, in the demo, is a verification system built to assist with purchases and ensure a secure and legitimate transaction after an agentic shopping session.

To register, users pick three sounds they can consistently repeat, such as a hum, a vowel sound held, a word, or a phrase. Each one of these sounds would be recorded on five occasions through alternating rounds in order to let the system understand their natural voice within these scenarios. With the setup complete, all the user needs to do is repeat two of the three sounds they select at checkout or during a bank verification call to be verified.

After every successful verification, the system also learns based on the attempt that passed the threshold. This way, as your voice patterns change, the system will gradually learn as well, lessening the chances of having to redeploy new sound choices constantly.

How we built it

The demo store and the bank are each built with FastAPI in Python, with an SQLite database connected to each one. The voice identification algorithm is built with a SpeechBrain ECAPA-TDNN speaker model, setting up thresholds based on the spread of the person’s own sound recordings. Payments are authorized in Stripe test mode along with the sandbox provider from Visa CyberSource. The agentic shopping experience uses GPT-4o to interpret requests and gather required ingredients for border statements, such as “Buy all the ingredients necessary for a tuna salad.”

Challenges we ran into

Recording the audio tracks back-to-back meant very little variety in the data collected. This made the model inaccurate and caused our initial testing to fail. We decided to pivot and switch to a rotating system where these audios would be recorded on a cycle with gaps in between to allow for a greater variety in testing data and more accuracy in the identity authentication process.

Another issue we faced was simply the accuracy of the authentication algorithm as well. We initially compared all test recordings to the entire set, but this resulted in a high percentage of false rejections at 72.9% for dysarthric speakers. By removing the single data point that was being tested from the entire set it was being compared to, this almost halved the percentage of false rejections based on the same set of data at 36.3%, leading to a substantial increase in the performance of the system. With this change, false accepts only rose from 0.00% to 0.83%.

Accomplishments that we're proud of

We’re proud of developing a solution that, while currently only integrated into a demo, can easily become a widespread verification system for banks and other programs. The ability for the algorithm to also learn and adapt over time is also something we are really proud of, as nothing will be a cookie-cutter checking system. Finally, building the demo website, which is also very accessible, allowing both typed commands and spoken ones, to provide a simplified, easier shopping experience rather than struggling to search for the right items.

What we learned

Entering this project, we had little to no experience working with audio. We learned how to take audio and analyze it in different ways. For example, sometimes we used speech-to-text to analyze audio and transcribe the audio to a rough sketch of what the user may have wanted, as well as using audio for authentication.

What's next for Timbre

Given the time constraint, we want to expand Timbre further into user testing, apply for greater data access, and continue to improve the algorithm and methodology of testing the voice authentication. Being able to integrate this system into real bank calls and verification systems beyond finance would be the ultimate goal of Timbre, as it would support such a widespread community, not just for those who live with dysarthria, but for greater groups that struggle with this current verification system.

Built With

Share this project:

Updates

Submission history