Inspiration
Most AAC devices for people who can't speak still come down to spelling one letter at a time, through an eye-tracker or a switch. It works for some, but not everyone.
Eye-tracking is the standard tool for someone with advanced ALS who is fully paralyzed but still fully aware, called a "locked-in state." It works because eye muscles are usually the last thing ALS affects. But that's not guaranteed. Eye control can decline too, and losing precision is one of the most common reasons people give up on their eye-tracker. Pushed far enough, some patients lose eye movement completely. At that point, an eye-tracker is just a screen nobody can look at on purpose, and a brain-computer interface becomes the only channel left.
The problem is that the channel disappears completely, no matter how fast it used to be. So we built something that doesn't rely on the eyes at all: a system that reads intentional EEG signals in the brain controlled by the patient. It's paired with a personal memory graph, so the person isn't just spelling, they're choosing from meaning, grounded in who they are and who they're talking to.
What it does
CogniVoice pairs an EEG headset with a personal memory graph and an LLM pipeline. It transcribes whoever's talking to the user, identifies them, and pulls relevant memories. An LLM proposes three short intents plus Spell and Cancel options. The user picks one by step-scanning: a trained mental command moves a highlight across tiles, a jaw clench selects.
The chosen intent becomes three full candidate sentences, grounded in real facts and personalized to the listener. Using Eleven Labs, Picking one speaks it aloud in the users voice, and the exchange updates the memory graph, visualized live as a growing 3D graph.
Tracks We Built For
Microsoft — This isn't a chatbot with a nicer UI — it's a step-scanning interface driven directly by an EEG headset, solving something still genuinely inaccessible today: real, non-invasive, two-way conversation for someone who's lost speech and reliable movement. The demo does not demonstrate someone querying an assistant, it demonstrates someone who is finally able to conversate with their loved ones and family again.
ElevenLabs — Every sentence comes back in the user's own cloned voice, built from one short consented voice sample, not just a generic text-to-speech voice. Talking to their spouse, kids, or grandkids still sounds like them, with local and browser fallbacks if the network drops.
Gemini — Runs the entire language layer: turning a chosen intent into three grounded, personalized sentences, figuring out who's actually speaking to the user, and building the memory graph itself from a life story. It's the difference between the same intent becoming something like "love you too, mija" for a daughter, and something more measured for a home nurse, generated fresh, every turn.
Tiger Data — We built the memory graph, every conversation, and all scan telemetry into one Postgres-based store, with a durable local backup so a dropped connection mid-conversation can't lose or duplicate a memory someone just shared. It's the foundation for a system that actually remembers someone, turn after turn.
How we built it
Three processes communicate in real time: one reads the EEG headset and turns raw signal into two trigger events; a backend runs the conversation logic, memory retrieval, and generation; a frontend shows a fullscreen pilot view and an operator dashboard.
The memory graph is retrieved in stages — vector search, graph expansion, then a boost toward whoever's currently speaking, so the same intent produces a different sentence for different listeners. Generation runs through an LLM with automatic fallbacks, so a turn can complete even if a provider is down. Speech uses a cloned voice, so conversations with family still sound like the user, not a machine. We're also migrating storage to a more durable database, with a local backup so a dropped connection can't lose or duplicate a memory.
Challenges we ran into
One of our biggest challenges was figuring out which electrodes to use to reliably detect the gestures we wanted to track. We had to experiment with different electrode placements and gestures to determine which ones produced the clearest and most consistent signals. Getting the headset to distinguish between different intended movements was especially challenging, but experimenting with the setup helped us find an approach that worked for our prototype.
Accomplishments that we're proud of
The moment we're proudest of: the same chosen intent turning into a completely different sentence depending on who's listening, warmer for a daughter, more matter-of-fact for a nurse.
We're proud we built a memory system that's durable by design a dropped connection mid-conversation can't cost someone a fact they just shared. And we're proud the whole system stays honest with the people it's built for: every fallback and synthetic signal gets disclosed, never hidden, and it speaks in the user's own voice, so a conversation with family still sounds like them.
What's next for CogniVoice
One of the next steps we have been considering is partnering with faculty in brain-computer interfaces or assistive technology to properly validate this and publish it, instead of it staying a hackathon demo.
From there, we want to move to testing this on real patients under a proper clinical protocol and look at extending this beyond ALS to any condition that leaves someone's mind intact but takes away speech and movement, like certain strokes or spinal cord injuries.
Built With
- 3d-force-graph
- asyncpg
- deepgram
- elevenlabs
- emotiv
- emotiv-cortex
- fastapi
- faster-whisper
- gemini
- numpy
- pgvector
- piper
- postgresql
- pydantic
- python
- react
- sentence-transformers
- three.js
- tiger-data
- timescaledb
- typescript
- uvicorn
- vite
- websockets
- zeromq



Log in or sign up for Devpost to join the conversation.