Inspiration

My grandmother is 100 years old. She lives with my mom and my sisters, and we take turns caring for her. Each of us handles something different. One keeps track of her medications. One knows what calms her down. One is awake at night when she gets confused. Nobody holds the whole thing.

The gaps are where things go wrong, and they're never dramatic. She drinks a little less for three days. She's a little more confused at night. By the time somebody says it out loud in the family chat, we've lost a week.

The chat isn't a record. You can't search it, you can't compare Tuesday to last Tuesday, and whoever wasn't there never finds out.

Then I started thinking about every other household and volunteer circle doing the same thing with nothing behind it. Whole neighborhoods run on this. Rotating people, shared memory, nothing written down.

it does

Baton is the shared memory for a care circle: a family rotation, a parish group, a neighborhood network.

There's a live web app you can try at batoncare.org, no install. It has two views. The caregiver app, where you record a voice note and get a brief on the person before you walk in. And the coordinator portal, where the day's alerts, each person's history, and the rotation live. It runs in Spanish and English, switchable per user, because the people doing this work don't all read the same language.

After each shift, a caregiver sends a voice note of 20 to 60 seconds. No forms. Baton transcribes it and pulls out what happened across a closed vocabulary of ids, sleep, mood, orientation, mobility,medication, pain, skin, social, household, incident, task), then updates a living record for that person.

From there it does three things.

Handoff sheet. One page for whoever takeemergency room: current medications,allergies, what calms her, how she communicates, who to call, who decides.

Pattern watch. Every day it compares each person against their own history. Not thresholds. Their own normal.

Coverage. It notices unlogged visits andks the group to fill them.

One rule sits underneath all of it: interrupt a human only when a human has to decide. Baton never diagnoses. It never contacts the person being cared for. It nevecoordinator's back, and it never resolvesits own alerts.

The demo is loaded with a real-shaped care circle, 24 elders and 14 caregivers across 437 visits, so you can watch the detection fire instead of reading about it.

How we built it

Python, on the Strands Agents SDK, against A one repo.

Layer What lives there
core/ Data models, the 13-category vocabulary, per-person baseline, change detection, storage. No HTTP, no model calls, fully unit tested.
agents/ Strands agents. intake runs on Claude Haiku 4.5, which is cheap and gets called once per voice note.
watch runs on Claude Opus 5, which is expeeir tools call core/ directly, never overHTTP.
api/ FastAPI on AWS Lambda behind API Gateway via Mangum. The only thing clients ever talk to.
web/ A bilingual progressive web app. five for the coordinator. Nine screens, notone more. Built on a design system with a contrast-verified palette, so an alert is readable on a cheap phone at 6am.

The detection rule

This is the part that has to be right, so it's plain arithmetic with tests. No model anywhere near it.

A person's baseline is how often each catego a 21-day window that stops short of theperiod being judged:

$$ b_p(c) \;=\; \frac{\left|{\, d \in W : c \text{ ran worse on } d \,}\right|}{\left|{\, d \in W : c \text{ was observed on } d \,}\right|}, \qquad W = [\,t - 31,\; t - 10\,) $$

A category deviates on day \(d\) when it ran worse and that isn't usual for this person:

$$ D_p(d) \;=\; {\, c \;:\; c \text{ worse on } d \;\wedge\; b_p(c) \le 0.34 \,} $$

A pattern break, which is the only thing that reaches a human, needs two or more categories deviating on two or more consecutive visited days:

$$ \exists \; k \ge 2 \;\text{ consecutive visited days } d_1, \dots, d_k \quad\text{such that}\quad \left| D_p(d_i) \right| \ge 2 \;\; \forall i $$

Someone who sleeps badly most weeks has a high \(b_p(\text{sleep})\), so a bad night isn't news about them. One bad day isn't a signal, and one bad category isnsited neither extends a streak nor breaksone, because absence of data shouldn't count as good news.

Falls, bleeding and chest pain skip all of it and escalate immediately.

e the work is split

This is the technical argument of the project.

Deciding whether there's an alert is code. core.baseline is deterministic and tested. A model doesn't get to decide that someone is fine.

Sweeping all the people is code. The sweep runs before the agent is even constructed.

Writing is the model. Turning a deviation and four quotes into two sentences a tired coordinator can read at 6am, in her own language.

Data is stored language-neutral, along with the source language of the original audio. Language resolves at the edge, either from Accept-Language or a ?lang= param. UI strings come from locale files and model output is generated in that locale. Nothing gets translated in the data layer.

Everything runs on AWS. The agents deploy toith the daily watch triggered by EventBridgeScheduler. The API runs on Lambda behind API Gateway, and the web app is served from Amplify behind CloudFront, with TLS from ACM and DNS on Route 53.

Challenges we ran into

Getting Bedrock to answer at all. Anthroree separate things, and the error messagemes one of them. You need the Anthropic use-case form submitted, a valid payment method registered specifically for AWS Marketplace, which is not the same one that pays your AWS bill, and the Marketplace agreement accepted per model through create-foundation-model-agreement. Missing the third produced an INVALID_PAYMENT_INSTRUMENT error that pointed at the second. It also failed intermittently, which was worse. Three calls in a row would succeed, then three would fail two minutes later, so every wrong fix looked like it had worked.

A rolling window that ate the thing it was looking for. The first baseline used a plain trailing 21 days. It never fired on slow decline and it took me a while to see why. The window contained the change. Eight days of bad sleep quietly became "how he sleeps," and the drifbaseline now stops 10 days short of theperiod it's judging. That one fix is the difference between catching a slow decline and never seeing one.

Asking an agent to be a for loop. The first watch agent got a list of 24 people and one tool per person. It would announce "I'll now review all 24" and then end its turn without reviewing any of them. An agent you ask to be a for will eventually forget to be one. The sweep moved into code.

A token budget that failed silently. Opus 5 has thinking on by default, and the max_tokens budget covers thinking and the response together. At 8,000 "I'll sweep now," and run out. No error,just a plausible answer that did nothing. It needed 32,000. If an agent on Opus 5 looks lazy, suspect the budget before the prompt.

A prompt obeyed too literally. The instruction "Do not describe what you are about to do, do it" made the model say nothing and do nothing. It complied with

Keeping a health agent out of diagnosis. Asking the prompt nicely isn't a guardrail. The model would slide from "drinking less, more confused" into naming a condition.

Accomplishments that we're proud of

The prompt asks, the code verifies. A runtime check validates every note the model writes against a list of forbidden clinical terms, in both Spanish and English. If a note names a condition, it gets rejected and sent back to be rewritten. The guardrail doesn't depend oood.

A model having a bad day can't turn a finding into silence. The sweep happens before the agent runs, so if the model writes nothing or skips someone, the deterministic text from core.baseline is used and the alert still goes out. The worst case is a badly worded alert,

Stable across languages. Four runs out of four, about 13 seconds for 24 people, 2 alerts raised, with the model quoting each caregiver in the language they actually spoke.

Safety lives in one function. Every recommendation is phrased in a single place that must never name a condition and never tell anyone what to do to a patient. It also stays free of pronouns, since the record doesn't store anyone's gender and has no business guessing it from a name.

**A seed dataset that keeps the demo honest.437 visits, generated deterministically, sothe alerts you see are found rather than planted at render time.

It's cheap enough to give away. The wholcent per voice note and twenty-two cents aday to watch two dozen people. The groups that need this most are the ones with no budget, and that shaped the architecture.

It's something you can actually open. Not a notebook and a README. A deployed bilingual web app with both sides of the workflow, running against the same agents described here. Record a note as a caregiver, then switch to the coordinator view and see what the system did

What we learned

The failures above turned out to be the same mistake wearing different clothes: giving the model a job that belonged to code. Iterating over a list. Deciding whether a number crossed a line. Guaranteeing a rule was followed. Every time I moved one of those into Python and left theds judgment and language, the system gotcheaper and more reliable at the same time.

A closed vocabulary is what makes any of this work. Free-text notes read beautifully and can't be compared. "She seemed a bit off" on Tuesday and "quieter than usual" on Thursday are the same observation, and no system can tell. Thirteen fixed categories is what turns a pile of voice notes into a baseline.

Concentrate the safety posture instead of scattering it. One testable function beats the same rule repeated across six prompts and enforced by none of them.

And verify model access with five consecutive calls, not one. An intermittent failure is worse than a total one, because it rewards you for the wrong fix.

What's next for Baton

  • [ ] Coverage agent. Unlogged visits and holes in the rotation, asking the group to fill a gap before it becomes a missed day.
  • [ ] Prescription refills computed from the prescription date, surfaced before the pharmacy run becomes an emergency.
  • [ ] A real messaging channel, WhatsApp or SMS, so a caregiver never has to open an app to send a voice note.
  • [ ] Multiple circles per coordinator, en one family and a neighborhood.
  • [ ] Long-term memory per person, outliving the 21-day window.
  • [ ] Native apps, so a voice note is twnstead of a browser tab.
  • [ ] Offline capture, because the caregivers who need this most are standing in a kitchen with one bar of signal.
  • [ ] A funding path for the circles that need one. Most of them run on somebody's own money.

Built With

  • amazon-api-gateway
  • amazon-bedrock
  • amazon-cloudfront
  • amazon-eventbridge
  • amazon-route-53
  • amazon-transcribe
  • amazon-web-services
  • aws-amplify
  • aws-lambda
  • bedrock-agentcore
  • boto3
  • claude
  • claude-haiku-4.5
  • claude-opus-5
  • fastapi
  • mangum
  • pwa
  • pydantic
  • pytest
  • python
  • ruff
  • strands-agents
Share this project:

Updates

Submission history