Inspiration

William’s younger brother has autism, and everyday conversations can sometimes be difficult for him. He is one of the kindest, most caring people you could meet. But he might miss a question directed at him or not pick up on a subtle hint that a friend needs to leave. To others, those small moments can be mistaken for disinterest or a lack of care, when the reality could not be further from the truth. He cares deeply about the people around him, and it hurts us to know that misunderstandings like these can make it harder for neurodivergent people to form and maintain the connections they value. For William, this became more than something he saw at home. As a CS and Neuroscience student conducting clinical psychology research with autistic people, he kept hearing versions of the same story in the many conversations he had while getting to know them: social connection is not always difficult because someone does not want to connect, but because knowing what to say or do in the moment can be difficult. And sometimes, by the time you realize what you wanted to say, the moment, and the opportunity to connect, has already passed.

This is personal for all of us, and maybe even for some of you. Abishek, Colin, and Ary also have distant relatives with autism, or are neurodivergent themselves. So we wanted to build something that directly helps them connect in the moment rather than just outcasting them to special education classes or therapy sessions. We hope that by using our technology, we can help bring people closer together.

What it does

Situational Awareness runs on Meta Ray-Ban Display glasses with a companion iPhone app. It listens to the conversation and shows one short cue on the lens at the right moment to help neurodivergent people follow along and be part of the conversation.

  • While the conversation is flowing, it suggests something relevant you could add, so you can jump in with your own thoughts instead of just listening.
  • When someone asks you a question, it helps you respond, like "Answer, then ask how their weekend was."
  • When someone finishes talking and you don't know how to respond, it suggests the next line.
  • When someone uses an indirect phrase like "I should let you get back to it," it explains what that often means.
  • When your own words come out bluntly, it suggests a way to recover.
  • When you want help instantly, you can manually ask for a cue by tapping your fingers together.
  • When the conversation ends, it gives you a short recap of what you talked about, so you can remember details and follow up later.
  • The recap also explains slang or phrases from the conversation, so you understand their style next time.
  • The whole time, it stays discreet, so no one else can tell.

We know neurodivergent people tend to be overstimulated and desire peace. So for the rest of the time, the lens stays quiet to help you focus. Live captions for each speaker stay on the phone, so the glasses display never gets cluttered. It can also remember friends you choose to add through facial recognition, so cues can follow up on things they care about.

How we built it

  • The iPhone app is built in Swift and SwiftUI, using Meta's Wearables Device Access Toolkit to stream the glasses camera and audio and to draw on the lens.
  • Meta's Muse Voice gives us live transcription with speaker labels to detect different voices.
  • Meta's Muse Spark reads the conversation and decides whether a cue is needed. It also handles camera scene checks.
  • Cues fire at moments instead of being timer-based. The app watches for questions, pauses, and indirect phrases, and it starts a check while the other person is still finishing their sentence (or mid-conversation to help you join in when you are feeling left out).
  • The app detects the wearer automatically. The glasses mic sits next to your mouth, so it learns that the loudest voice is you and never responds to your own words.
  • A small Node.js server holds all API keys, so nothing secret lives on the phone.
  • On-device FaceNet matches friends the wearer enrolled, with no face data sent to any model.

Challenges we ran into

  • Our initial issue was speed. Our first version checked the conversation every 10 seconds, and the model took about 10 more to answer. Cues showed up 20 seconds late. We rebuilt it to react to moments in the conversation and to send the model only the recent text, so cues now arrive while there is still room to use them.
  • Cues used to be useless. Early on, the model always tried to say something, so it gave filler statements like "stay attentive." We rewrote the rules so it only speaks up when it has something specific to add.
  • The model had no idea who was talking. In a group, several people talk over each other. So we needed live captions that separate each speaker, and cues that respond to the right person instead of mixing everyone's words together.
  • The Meta Ray-Ban Display is physically limited to a small area in the right lens. Visually showing captions and cues together is overstimulating, especially for neurodivergent people, so we cut the lens down to one good cue that becomes your way into the group conversation.

Accomplishments that we're proud of

  • A working end-to-end system on real Meta Ray-Ban Display glasses.
  • Cues that arrive during the conversation, instead of about 20 seconds late.
  • Cues that add to the conversation so naturally that no one can tell you are getting help.
  • Over 300 automated tests across the app and server.

What we learned

In a live conversation, the timing of certain phrases matters more than intelligence. For someone who struggles to know what to say, the best response is one that is readily available while there is still room to use it without being awkward. A lot of the hard problems we faced were not about AI at all. Knowing who is speaking and fitting everything into a crammed UI took as much work as the model, and getting a cue to show up at the right moment took as much work as making it intelligent.

What's next for Situational Awareness

  • Test with autistic users and their families, starting with the community William works with, and design the cues with them rather than for them.
  • Hear tone of voice. Our model reads words, not tone. An audio native model could catch sarcasm or a flat "fine."
  • Faster end to end. The model is fast now. Next is shaving time off transcription.
  • Stronger privacy controls for friend recognition, so everyone in a conversation can trust it.

Built With

Share this project:

Updates

Submission history