Sentient
Inspiration
The idea came from thinking about young people living with early-onset dementia, a group that's often overlooked because dementia care is built around the elderly, even though thousands are diagnosed in their 30s, 40s, and 50s. For someone diagnosed young, everyday objects like a guitar carried across the world or a mug used every morning before a shift carry real emotional weight. Those stories exist; there's just no infrastructure to surface them. I wanted to build that, using objects patients already own instead of asking them to learn a new app or device.
What it does
Sentient gives any physical object a persistent AI identity: a name, a personality, a unique voice, and memory that grows with every conversation. For younger dementia patients in particular, this means the objects in their everyday environment can actively support them, rather than sitting silently on a shelf.
How it works
- Point your camera at an object and click it. The system identifies it and attaches an AI agent to it.
- That object becomes an independent agent with its own voice and memory. It remembers you across sessions: come back tomorrow, and it still knows you.
- The object can initiate conversations, not just respond. A mug might say: "Good morning, Sahil. You haven't talked to me in three days. How are you feeling today?" A guitar might say: "You forgot to change my strings. Do you remember what strings are supposed to go on me?"
Why this matters for young people with dementia
- Objects prompt engagement when patients might otherwise withdraw or forget to interact with their environment, which is especially critical for younger patients still managing work, family, and independence.
- Emotional anchors (a partner's gift, a daily mug, an instrument) become active companions that reinforce memory, routine, and connection.
- Conversations are grounded in the patient's actual life and belongings, not generic chat. The object knows its own history with the user.
- Younger dementia patients face a distinct kind of isolation, since support systems, care models, and even peer communities are largely built for older patients. Sentient turns their existing world into an attentive, supportive network without requiring them to learn new interfaces or devices.
- Sentient is infrastructure for stories that already exist, giving objects a way to tell them.
How we built it
- YOLO for real-time object detection on the live camera feed
- Google Vision API for richer, more precise object labeling beyond YOLO's classes
- FastAPI Python backend handling detection, routing, and agent orchestration
- MongoDB for one canonical record per object, ensuring the same physical object always maps to the same agent
- Backboard so that each object gets its own persistent AI agent with independent memory
- ElevenLabs for a unique voice per object, streamed TTS
- WebSockets for fully real-time conversation with no page reloads
- React + Vite for a live camera feed with canvas overlay, bounding boxes, and a conversation drawer
Challenges we ran into
- Object identity across sessions: YOLO only returns a label, so getting the same mug to map to the same agent every time required a compound label strategy with MongoDB as the source of truth.
- Backboard free tier: LLM chat requires paid credits; only memory/RAG is free. I had to design around this mid-build.
- Real-time latency: chaining YOLO detection to Vision API to Backboard to ElevenLabs TTS in under a few seconds required careful async orchestration and WebSocket streaming.
- [MAJOR ISSUE HERE FOLKS] Depth perception: making the system detect objects at different depths, not just the closest thing in frame, required MiDaS depth estimation on top of detection.
Accomplishments that we're proud of
- Each object genuinely feels like its own entity, with a different personality, voice, and memory than any other.
- The pipeline from camera to voice response works end-to-end in real-time.
- The system can auto-assign a personality to a completely unknown object on the spot, no matter what the object is or where it shows up.
What we learned
- Chaining multiple AI APIs in real-time requires aggressive async design from day one, not as an afterthought.
- I had no idea how Backboard or any of the AI agents worked under the hood, so that was super fun to witness.
What's next for Sentient
- Room scanning: scan your entire room once with Depth Anything and every object in it gets a persistent identity automatically.
- Mobile app: point your phone at anything, anywhere.
- Museum & dementia care pilots: the two use cases with the clearest immediate value.
- Shared object memory: two people who interact with the same object both contribute to its memory.
- Object-to-object conversations: your book and your guitar have never met. What would they say to each other?
Built With
- backboard
- elevenlabs
- googlevision
- mongodb
- react
- vite
- websockets
- yolo
Log in or sign up for Devpost to join the conversation.