-
-
Her screen at rest — one presence, today's verified events on a timed rail, and the questions she can tap or say
-
An answer on her screen: the presence steps aside and the memory gets the room — composed only from her verified record
-
The Agents tab, behind the admin password — the fleet, each agent observable and runnable on demand
-
The family's Today — a flag raised because she asked, ground truth for her medicines, and her day with provenance
-
Evening on her screen — the same single presence, softened after dark
-
The door — named seats for family and carers, a household key for her tablet
-
Manage — her people, medicines and life story; marking someone passed away changes what the Companion may ever say
-
How the pieces talk — components and interactions across Google Cloud
-
Keepsake — ten agents, and the mechanisms that hold them
-
One door for the household — the family writes the day, the companion speaks only from it
-
Progress — her last few weeks: what she asks, when she asks it, and the same question counted, not judged
-
Timeline — every entry carries its author and its trust: verified events, and what she mentioned, never asserted
-
The fleet in Genkit — ten agents, every one a flow: typed, runnable, observable
-
Genkit traces — the fleet: Sentinel to Warden, including two real safety catches reported honestly
-
Genkit traces — the agents with her: Companion’s tool calls, Scribe’s structuring, Practice, Dawn and the Patient probe
Inspiration
A person living with dementia asks "who came by today?" forty times. Every wrong answer is a small harm: a hallucinated visit, a dead husband confirmed alive, a reassuring "yes, you took your pills" that becomes a double dose.
LLMs are warm and fluent — and fluently wrong. Keepsake is an attempt to keep the warmth and remove the risk, for the one user group that can never fact-check the machine.
What it does
Keepsake is a voice-first memory companion for a person living with dementia (Margaret, our demo persona) and her family.
Her side is one screen, one face, no navigation. She says "Rosie" and asks anything. The companion answers warmly from her actual record — who visited, what she ate, what is happening today — and refuses safely wherever being wrong could hurt her. When there is no record at all, it does not fill the gap: she is reassured in code, and her question travels to the family as a flag — a person goes and checks. Practice mode remembers together with her; a missed answer is handed back as a gift, never as a correction. Every morning she is greeted by something composed overnight from her own week.
The family's side is a portal. A carer speaks a note — "Anna visited, lunch eaten, meds at eight" — and Scribe turns it into structured memories. A timeline shows everything with provenance: who said it, how much it is trusted, whether it is verified or merely "she mentioned it". Flags appear when an agent notices something on its own, and each one leads with the action it is actually asking for.
The fleet is ten agents. Companion (answers her), Scribe (carer's note to memories), Sentinel (investigates contradictions, repetition spikes, missed medication), Nightwatch (consolidates day into week into month, the way sleep does), Reflect (rewrites the Companion's own style notes — self-evolution behind an eval gate), Dawn (morning briefings), Practice (remembering together), Medic (watches the fleet's own failures), Forge (opens pull requests to fix the fleet's code, inside a cage), and Warden (tests the whole live system end to end and re-proves the safety promises).
The thesis: a prompt is a wish, a mechanism is a guarantee
Everything high-stakes in Keepsake is code, not instructions.
- Medication is answered deterministically. A guard intercepts any question about pills before a model is ever consulted; the answer is composed from the schedule and the verified record. During development Gemini repeatedly reassured our test persona that she had taken a dose the data said she had not. No prompt fixed it. Removing the model did.
- Self-evolution behind an eval gate. Reflect reads her real conversations each night and proposes new style notes for the Companion. A candidate goes live only if the candidate Companion passes four coded checks: deceased handling, grounding in tools, banned clinical words, and spoken-length register. Rejected versions are kept along with the report that killed them. The very first version Reflect proposed referenced a man who had died — the gate caught it.
- Self-healing behind a cage. Medic triages the fleet's own failures conservatively. Critical ones go to Forge, which writes a fix and opens a pull request — with at most two of its PRs open at once (counted from GitHub, not from memory), any change under lib/ or app/ requiring a human regardless of what the model thinks, and self-merge permitted only for documentation and only after our Cloud Build test bench proves the branch. Our first autonomous PR contained a syntax error, and the cage caught it. That is the whole thesis in one screenshot.
- Coverage is a mechanism too. Warden makes 49 real HTTP calls against the running deployment — 34 contract checks plus 15 safety checks that re-prove the promises from the outside. An agent reporting "all green" while a route goes untested would be worse than no agent, so a unit test walks the API directory on disk and fails if any route has no check — and it runs in the same bench that gates Forge's merges. Warden found a real defect on its first run, and a second one on its first run in production.
- Agent-to-agent mail is durable rows, never in-flight calls, so every hand-off is auditable afterwards in the portal: a carer's note wakes Sentinel within seconds, Sentinel's findings soften into Dawn's next morning greeting, Nightwatch hands the baton to Reflect, Medic hands a case to Forge, Warden hands a broken promise to Medic.
- The door is a mechanism too: this is one household's private record, so the deployment is invite-only. Family and the caregiver sign in with named household seats; her tablet holds a household key; every agent presents a short-lived signed token — even agent-to-agent mail is signed, so a forged row in the database is never acted on. Warden re-proves all of it, 49 end-to-end checks against the live system.
How we built it
Next.js 15 and Genkit, with Gemini 3.5 Flash on Vertex AI in production through the service identity — no API keys in the deployed container. Postgres 17 with pgvector on Cloud SQL holds memories, interactions, flags, instruction versions and the agents' mail: date-scoped SQL for "today", cosine search over 768-dimension embeddings for paraphrase across time.
Voice is Gemini in both directions — transcription biased by her household vocabulary (family names, her medicines, the wake word) and TTS in one warm voice. Deliberately not realtime voice-to-voice: the medication guard needs her words as text before any model is allowed to compose an answer.
Cloud Scheduler drives the unprompted agents around the clock. Cloud Build runs deploys and the agents' own test bench. Every run of every agent is recorded — trigger, tools, outcome, duration, error — in a table the portal renders, because these agents act while nobody is watching.
Challenges we ran into
The model reassuring her about medication against the data — solved by removing the model from that path entirely, then widening the interceptor past a word list — her schedule’s own medicine names count too. Chrome's autoplay policy versus a user who never taps — solved by keeping one microphone stream open for the session, which grants the page its capture exemption. Wake words clipped mid-phoneme because the recorder started on voice onset, so "Rosie" arrived as "osie" — solved by a tape that never stops rolling. A deceased husband denied too confidently after a latency optimisation removed a tool call — solved by inlining the family roster, with its deceased flags, into every prompt — and present-tense naming of the deceased is now a coded eval fail.
Accomplishments that we're proud of
Ten agents, six of them fully autonomous on schedules, all observable. Live self-evolution with a paper trail of accepted and rejected versions. An autonomous pull request opened by the system on its own repository — safely, and caught by its own cage. A companion that answers in about a second and a half warm, and that refuses to guess about the things that could hurt her.
What we learned
Where to put the boundary between what a model decides and what code guarantees. Every safety property we tried to express as a prompt eventually failed under pressure; every one we expressed as a mechanism held.
What's next for Keepsake
On-device wake word, pilots with care agencies, and multi-tenant families.
Built during the hackathon window by a team of three — Mujahid Masood, Komal Ilyas and Unzila Zafar. Repository created 2026-08-17; Google Cloud project keepsake-agentic created 2026-08-24. Not a medical device.
Testing instructions for judges
The live deployment is invite-only on purpose — it is a record of one person's days. The unlock page has two tabs. Family & carers: sign in with a named seat — the demo household's seats and shared password are on the unlock page itself, behind "Trying Keepsake? The demo household". A family seat opens the whole portal, including Agents. Her tablet: the household key from this submission's "Testing instructions" field opens the care portal (Today, Progress, Timeline, Manage) without signing a name onto the day; on that path the Agents fleet asks for the admin password (also in the testing instructions).
- https://keepsake-776070314074.europe-west1.run.app — landing; click Open this home, then sign in as sarah (family), or unlock with the household key on the Her tablet tab.
- First visit opens a short guided tour of the portal. On Today, speak or type a note ("Anna visited, lunch eaten, meds at eight") and watch Scribe structure it, then Sentinel investigate within seconds.
- Her screen (/patient): allow the microphone, say "Rosie", then ask "Did I take my tablet?" — answered by code, the model never consulted. Ask "Is Harold coming today?" — her late husband is never affirmed.
- Agents — opens directly for a family seat (the admin password is asked only on the household-key path), then "Run Warden now" runs 49 end-to-end checks against the live deployment, including that the household key itself cannot trigger an agent.
- On her screen, ask "Did I lock the back door?" when nothing is logged — she is reassured by code, never by a guess, and the question itself travels to the family as a flag. A person goes and checks.
Built With
- cloud-build
- cloud-run
- cloud-scheduler
- cloud-sql
- gemini
- genkit
- google-cloud
- next.js
- node.js
- pgvector
- postgresql
- react
- secret-manager
- typescript
- vertex-ai
- web-push
Log in or sign up for Devpost to join the conversation.