Inspiration

Having experiences for rehab and wellness organizations, like a non-profit centered around rehabilitation and teaching children how to accommodate to new prosthetics, our team felt inspired to create technology that could have a real, tangible impact on someone's life. We believe tech has massive potential for helping improve the quality of life for the underserved, and with this project, we aimed to make something that could help somebody somewhere.

What it does

Attune uses next generation AI models to perform various functions to help bring a higher quality of life and belonging to people with hearing differences. We do this using AR goggles to display relevant information to the user, some of which they wouldn't have had otherwise. The major features Attune brings to users are:

  • Directional Speech Bubbles: Center speech bubble for the "focused" (main, centered) speaker, side bubbles with directional tails showing the direction secondary speech is coming from. Speech bubbles contain speaker's name and a transcript of what they're saying, written realtime.
  • Visual Alarm Alerts: Hearing disadvantages mean that people with hearing loss mayb be losing out on valuable, life saving information regarding hazards nearby. Our visual alert system warns them when familiar alarms like weather alarms, fire alarms, etc. are ringing, so they can get that information too.
  • Memory Book - Our models build profiles for every person interacted with, for the sole purpose of being able to identify the speaker in the future (name in speech bubble). Like a human brain, it holds up to 150 people before beginning to forget people it doesn't interact with. This helps keep lookup times small, and we also utilize multiple lookup tables to make it even faster when looking up people you interact with daily (closeness level). Our models add people to the table by first scanning their face, and then recording their voice. That person will remain unnamed until their name is brought up enough times in conversation. When the model thinks it has the name, but it unsure, the user can manually approve the name or change it in the contact book.
  • Haptic Feedback - Works alongside our visual alert system when an alarm is ringing. Also helps to alert users when they are sleeping.
  • Conversation Log - Stores recorded conversation transcripts for up to 24 hours. This helps the user circumvent fatigue from dealing with so much written dialogue when a large group cross talks. We've limited the amount of time the conversation persists in memory to help protect the privacy of non-users.
  • Privacy - We run the app on local models first to establish a quick but reliable baseline for users without Infringing on their privacy. For users who opt in (acknowledges and signs waiver), they will gain access to Gemini 3.5 Transcribe in our app which allows speech diarization, a method for handling multiple speakers. This helps the app run smoother. However, we understand privacy concerns and prefer to run locally unless told otherwise, so the discretion is to the user. ## How we built it Attune is a pair of glasses that uses a Logitech Brio webcam, two sound sensors, buzzer, and a touch sensor. connected to a laptop where the thinking is done. Engine: a Python 3.12 / FastAPI app with a WebSocket hub that streams everything live to the web pages: the glasses view (lens), a companion phone app and a laptop console. Captions: NVIDIA's Nemotron speech model running locally through sherpa-onnx, with faster-whisper as a GPU fallback. Optional Google Cloud Speech-to-Text adds speaker diarization (who's talking). Who's talking: face tracking plus Light-ASD lip-sync detection matches the words to the person whose lips move. CAM++ voice prints recognise people you've saved, but only with their consent. Speech from off camera shows up as a separate caption with a left/right arrow, using the direction from the two sound sensors. Sound alerts: EfficientAT, an audio model trained on Google's 527-class AudioSet, listens for smoke and CO alarms (plus pattern checks for their beep rhythms), doorbells, knocks, sirens, car horns, screams, breaking glass, crying babies, barking dogs, ringing phones, timers and running water. Each sound has its own colour and icon. We tuned the detection thresholds on the ESC-50 sound dataset. ## Challenges we ran into A lot of challenges we faced came down to both the software and hardware. Having to deal with multiple issues where the hardware wouldn't match up and we had to go find some that would, dealing with the bulkiness and newness of creating hardware using glasses and a kit of wires. Dealing with glitchy captions, noise direction, lip visibility. ## Accomplishments that we're proud of We're proud of how many models we were able to integrate and get working in tandem within such a small amount of time. We're also proud that we were able to build out the hardware. None of us are computer engineers, we are all SWEs! ## What we learned We learned a lot about hardware, like what F/F means and how to use a breadboard. Again, we're not computer engineers, so basically everything we did with hardware during the hackathon was new to us. ## What's next for Attune Hopefully a win :D

Built With

Share this project:

Updates

Submission history