Inspiration

In the residential colony where I live, checking on elderly residents who live alone runs on a WhatsApp group and goodwill. Someone says "I'll visit Lakshmi aunty today" — and most days someone does. But when a visit gets missed, nothing notices. The family finds out days later. The failure isn't malice; it's diffusion: everyone assumed someone else had it covered.

I wanted to build the thing that notices.

What it does

  • Every morning, each resident is assigned a volunteer — with a fairness rule that looks at the last 7 days of history, so the same few people don't quietly end up doing every weekend.
  • The volunteer visits, then marks the check-in done with an OTP texted to the resident that morning. "Done" means somebody was actually there, not that a button got clicked.
  • If it isn't done, an AI agent escalates: volunteer → secretary → joint secretary → emergency contact — the chain RWAs actually use. At 60 minutes it reminds the volunteer; at 3 hours the secretary; at 5 hours the joint secretary; at 7 hours the emergency contact.

How I built it

The whole stack is AWS, deployed with CDK: Lambda, DynamoDB, API Gateway, Cognito, SNS, EventBridge, CloudWatch. The frontend is a React dashboard on Vercel.

The heart is the escalation agent, built with the Strands Agents SDK. It has four tools — get_pending_check_ins, get_person_by_role, send_notification, update_check_in_status — and an escalation policy written in plain English in its system prompt. There is no hardcoded if/else in the escalation path: the model reads today's pending check-ins, reasons over the policy, and decides which tool to call for each one. Every decision is appended to an escalationLog on the check-in record, so the full audit trail lives in DynamoDB.

The model runs on openai/gpt-oss-120b via Groq's OpenAI-compatible endpoint through Strands' OpenAIModel. Security is real authentication, not theater: Cognito with a PreSignUp Lambda that only permits sign-ups for secretary-registered emails, server-side role filtering, and a deliberately time-limited sign-in lockout.

Challenges I ran into

A model provider vanished mid-build. I originally targeted Bedrock + Claude, but a new-account activation wall pushed me to Groq — swapping providers was a ~10-line change, which is the best argument for Strands' model abstraction. Weeks later, Groq deprecated the model the agent was using. Nothing in my code was wrong; a third-party model just stopped existing. Lesson: any project pinned to a model ID should budget for deprecation.

The system failed successfully — for hours. A reasoning model once burned its entire token budget on internal "thinking" and returned an empty response with zero tool calls. The Lambda exited fine. Nothing errored, nothing escalated. A system whose entire job is noticing when something's wrong has to notice when it is wrong — so now CloudWatch alarms email me if the assignment job fails, if the agent errors twice in a row, or if it silently doesn't run at all.

Auth took longer than the agent. V1 was "pick your name from a list" — anyone could act as anyone. V2's PINs still let anyone claim an unclaimed identity. V3 accepted this was an authentication problem: real Cognito accounts, gated sign-up, and server-side enforcement so a volunteer can't fetch another volunteer's data even by calling the API directly.

What I learned

  • Write the policy in plain language before writing code. If you can't state the escalation rules in five bullets a non-engineer could verify, the prompt won't save you.
  • Boring tools are the feature. Mine read DynamoDB and write an audit log — which is exactly why the agent's behavior is inspectable.
  • Test auth as an attacker, not a user. Every real security fix came from trying to misuse the flow, not from design review.
  • Build failure alarms before you think you need them. You'd rather learn about a silent failure from an email than from an 8-hour-old dashboard.

What's next

Member removal and SMS delivery reliability are still rough edges, and it runs on synthetic data until a real pilot gets residents' consent. The core bet works: an agent that notices, decides, and leaves a receipt.

Built With

Share this project:

Updates

Submission history