Inspiration

Frank called his daughter three times last week asking who Marcus is. Marcus is his grandson. He's eight.

He doesn't forget because he doesn't care. He forgets because dementia takes memories before it takes people. And the people who love him remember everything.

Folkore started from watching this happen — not once, but in families everywhere. The gap between "the family remembers" and "the patient remembers" felt solvable. Not with a medical device or a care app — with voice, memory, and the people who matter. When Frank asks Alexa who Marcus is, Folkore doesn't pull from a general AI trained on the internet. It pulls from the specific things his daughter chose to tell Folkore — that Marcus is eight, that he wants to be a palaeontologist, that Frank calls him his little dinosaur guy. That distinction is the whole point.


What it does

Folkore is a voice-first memory companion for dementia patients and the families who care for them.

A family member sets up Folkore once — opening the web UI and adding memories: names, stories, places, moments that matter. Each memory is stored with who, what, when, and tags. Relationships are extracted automatically into a memory graph — Marcus links to his memories, Lake Tahoe links to its stories.

When Frank talks to Alexa and can't remember his grandson's name, Folkore retrieves the right memories and responds warmly: "That's Marcus — your grandson. He's eight and absolutely loves dinosaurs. He always brings a new dinosaur fact when he visits — and he has a big school project coming up he can't wait to tell you about."

Every morning, before Frank says a word, Alexa proactively surfaces a memory unprompted. He didn't ask. Alexa just remembered for him.

For Sarah, 400 miles away, the insight dashboard shows mood trends, confusion signals, and which memories Frank has been reaching for — without a phone call, without Frank having to explain. A weekly digest email summarizes the week in plain language: how he's been, what he's talked about, one thing worth adding.

No app. No screen. No training required. Just voice, memory, and the people who matter.


How I built it

Folkore is a TypeScript MCP server that exposes two tools to Alexa+: converse and morning_memory. Every request passes through a Supervisor Agent that classifies intent and routes to one of three specialist agents — each running a real agentic tool-use loop where the LLM decides which tools to call and in which order.

Four agents:

  • Supervisor — classifies every utterance and routes to the right specialist. No business logic lives here.
  • Conversation Agent — parent-facing. Runs hybrid three-layer retrieval, responds warmly. Confusion detection runs server-side via deterministic regex — the LLM never decides what counts as confusion.
  • Curation Agent — family-facing. Validates memory completeness, stores to DynamoDB and S3, extracts named entities into the graph, confirms back to the family.
  • Insight Agent — generates mood trends, confusion narratives, and the weekly digest email from live interaction data.

Three-layer hybrid retrieval:

  1. Graph traversal — named entity match → follow edges → collect memory IDs instantly
  2. DynamoDB keyword scoring — always fresh, catches memories added seconds ago
  3. Bedrock Knowledge Base — semantic vector search, best-effort augmentation

AWS stack: DynamoDB (6 tables), S3, Amazon Bedrock Knowledge Base, Bedrock Agent for KB sync management, Bedrock Mantle gateway for LLM inference via deepseek.v3.2.

Guardrails: PII scrubbing strips phone numbers, emails, SSNs, and card numbers from all memory fields before any write. Rate limiting at 20 req/min per IP prevents runaway spend.

Deployment: The server runs on AWS Lambda with a Function URL — a persistent public HTTPS endpoint judges can connect Alexa+ to directly, with no local setup required.

System prompts live in .txt files — version-controlled, diffable, readable without TypeScript. A clinician reviewing what Folkore says to a dementia patient can open a text file and read it. Changing a prompt never requires a code deploy.


Challenges I ran into

Getting retrieval right. A single query — "tell me about my grandson" — needs to find memories stored under different words, different tags, or linked through a graph relationship. Pure keyword search misses semantic distance. Pure vector search misses freshness. I built a three-layer pipeline that merges all three: graph traversal for entity-centric queries, DynamoDB for always-fresh results, Bedrock KB for semantic augmentation. The graph layer was the key unlock — "Marcus" resolves to a node, the node resolves to every memory about Marcus, before keyword scoring even runs.

KB sync concurrency. Bedrock KB ingestion is asynchronous. If a family adds two memories in quick succession, a naive implementation fires two StartIngestionJob calls and one hits ConflictException. I built a sync guard that checks ListIngestionJobs before every fire — if a job is already running, the file is already in S3 and will be picked up incrementally. DynamoDB keyword retrieval always covers freshness regardless of KB sync state.

PII in unexpected places. Families dictate contact details alongside memories — "Sarah's number is 555-..." — and phone numbers contain enough digits to trigger PII scrubbers unexpectedly. Even our own eval test markers (timestamp-based strings with long digit runs) hit the scrubber during testing. I tightened the scrubbing regex and switched all internal markers to alphanumeric-only strings.

React StrictMode double-fire. The morning memory trigger fired twice on load because StrictMode mounts, unmounts, and remounts components — a useState boot flag doesn't survive the cycle. I moved the boot flag to useRef, which persists across the StrictMode remount.

Keeping the demo honest about SNS. We wanted to show the weekly digest but didn't have time to wire EventBridge + SNS. Rather than hide the gap, we built a manual trigger and a clear UI badge: "Manual trigger · SNS delivery in Stage 3." The email is real — generated from Frank's actual data. The delivery mechanism is the staged part.

Stateless agent context across turns. The Curation Agent runs a fresh loop on every API call, so a follow-up reply ("he loves fishing") had no idea what the family member was asking about in the previous turn. Fixed by threading a capped conversation history (last 10 messages) from the UI through the request, injected into the agent loop before the current message. History resets automatically when the agent confirms a memory is saved


Accomplishments that I'm proud of

20/20 behavioral evals passing against live AWS. Five suites — tone, retrieval correctness, confusion detection, hallucination guard, and curation write/read — all passing against the real DynamoDB and Bedrock stack. No mocks. No LLM-as-judge. Pure string and regex assertions against deterministic system behavior.

Deterministic confusion detection. The LLM never decides what counts as confusion. A regex classifier on the utterance catches temporal signals ("what year is it"), person misidentification ("who is Dorothy again"), and place errors before the agent even runs. The result is a typed confusion_type enum in the log — auditable, consistent, not subject to model drift.

Prompts as first-class artifacts. Every agent's system prompt lives in a .txt file. Changing what Folkore says to a dementia patient is a file edit, not a code deploy. That's the kind of auditability a product like this needs in production.

Guaranteed logging. logInteraction runs unconditionally in the server wrapper after every turn — not as a tool call the LLM might skip. Logging persists even if the agent errors, hits max iterations, or returns an empty response. The insight dashboard's data is always complete.

91 unit tests with zero AWS dependencies. Confusion detection, agent loop behavior, PII scrubbing, rate limiting, and keyword scoring — all testable locally with no cloud setup.


What I learned

The hardest problem in a memory companion isn't storage — it's retrieval. Getting the right memory at the right moment, for a query the family didn't anticipate when they wrote it, is where naive approaches break down. The memory graph was the architectural decision that unlocked entity-centric retrieval: once "Marcus" is a node, every memory about Marcus is reachable through a two-hop traversal regardless of how the family worded it.

We also learned that honesty about staged features is better than hiding gaps. The weekly digest manual trigger and the "SNS in Stage 3" badge are more compelling to a judge than a faked automation. The email is real. The data is real. The trigger mechanism is the thing that's staged — and saying so explicitly is more trustworthy than pretending.

Deterministic guardrails matter more than LLM-level safeguards for a product in this space. Confusion detection by regex, PII scrubbing before every write, logging outside the agent loop — these aren't limitations, they're the right architecture for a system that touches vulnerable people.


What's next for Folkore

Stage 2 — Richer Memory

  • Voice-native curation — family members add memories by talking to Alexa directly, no web UI required
  • Photo memories — when Alexa opens photo ingestion to MCP tools, a vision model extracts names and context from family photos into voice-retrievable memories
  • Memory aging — surface recently added memories more often; rotate older ones back in on anniversaries or relevant dates

Stage 3 — Proactive Family Intelligence

  • Real EventBridge trigger — morning memory fires automatically every day per registered parent
  • Automated SNS digest — weekly email fires every Sunday without manual trigger
  • Confusion escalation — spike in confusion rate triggers immediate family notification, not just the weekly cycle
  • Memory gap detection — Insight Agent flags entities Frank mentions that have no memory attached: "Frank mentioned 'the Hendersons' twice this week — no memory exists for them"

Stage 4 — Production & Privacy

  • Multi-tenant isolation — per-family DynamoDB partitions, per-family S3 prefixes
  • HIPAA-aligned storage — encryption at rest and in transit, audit log of every memory access
  • Neptune migration — replace DynamoDB graph tables with Amazon Neptune for native multi-hop traversal as graphs grow past 500 nodes
  • Caregiver portal — read-only view for professional caregivers: confusion trends, memory coverage, conversation summaries — no raw transcript, ever

Built With

Share this project:

Updates

Submission history