Inspiration
Communication is one of the most fundamental parts of being human, and emotion is at the heart of it. When we started thinking about challenges people face, the experience of paralyzed individuals using AAC (Augmentative and Alternative Communication) devices stood out. These devices give people a voice, but the voice they produce is flat and robotic. "I love you" and "I need help" sound identical. Sarcasm, humor, warmth, and sadness all disappear. As college students at Georgia Tech, we wanted to tackle something that felt genuinely meaningful, and restoring emotional expression to people who have lost the ability to show it felt like exactly that.
What it does
Iris is a web application that gives paralyzed AAC users an emotionally expressive voice. A webcam reads the user's facial cues and eye movements to detect their emotional state. Based on what the conversation partner says (picked up by a microphone), an AI generates a few short suggested replies for the user to choose from using only their eyes. Once the user selects a reply and confirms a tone, the app speaks it aloud in a voice that actually sounds happy, sad, excited, or serious. A feedback loop lets the system learn each individual user's unique expressions over time, since paralysis affects facial muscles differently for everyone.
How we built it
We built Iris as a full-stack web application. The frontend is built with React and TypeScript, using MediaPipe Face Landmarker to process the webcam feed in the browser and extract facial landmarks and blendshape scores in real time. Eye gaze direction, blink patterns, and dwell selection let the user navigate without touching anything. The backend is built with Python and FastAPI, which handles calls to the Anthropic Claude API to generate contextual reply suggestions and manages the per-user emotion model built with scikit-learn. For emotional speech output, we integrated LiveKit with Cartesia TTS, which lets us control the emotion, speed, and volume of the synthesized voice for each message.
Challenges we ran into
Webcam-based eye tracking is nowhere near as precise as a dedicated eye tracker. Getting reliable gaze region detection without an expensive hardware setup required a lot of calibration work and careful threshold tuning. We also ran into the core problem with generic emotion detection: paralyzed users often cannot make typical facial expressions, so a model trained on healthy faces is nearly useless. Designing a per-person calibration flow that a caregiver could run in a few minutes, without requiring any technical knowledge, was harder than expected. Keeping the interface usable for someone who is exhausted from eye fatigue was another constant design constraint.
Accomplishments that we're proud of
We built a working end-to-end pipeline where a user can have a real conversation using only their eyes and have their replies spoken with genuine emotional tone. The interface is clean, navigable, and designed with the actual experience of a paralyzed user in mind, not just the technology behind it. We're also proud of the feedback loop, which treats each user as an individual and gets better the more it is used.
What we learned
We learned how much is lost in translation when emotion is stripped from speech, and how easy it is to underestimate that loss. We went deep into AAC research and found that tone of voice has been almost entirely absent from AAC development for decades despite users consistently asking for it. On the technical side, we learned how to work with real-time computer vision in the browser, build accessible interfaces under strict physical constraints, and chain together AI, voice, and vision pipelines into a single coherent experience.
What's next for Iris
We want to integrate Iris with dedicated eye tracking hardware like Tobii Dynavox devices, which are already in the hands of the people who need this most. Bridging Iris with those systems would mean users do not have to choose between their current setup and emotional expression. We also plan to expand the emotion vocabulary, add voice banking so users can speak in their own recorded voice, and test with real AAC users and caregivers. Longer term, we see Iris as a layer that sits on top of existing AAC software rather than replacing it, so millions of people can benefit without switching systems.
Built With
- anthropic
- claude
- computer-vision
- docker
- elevenlabs
- eye-tracking
- fastapi
- mediapipe
- meta-muse
- ml
- python
- react
- scikit-learn
- sqlite
- tailwindccs
- typescript
- vercel
- vite
- webgazer
Log in or sign up for Devpost to join the conversation.