Inspiration

In the 2021 heat dome, 619 people died in British Columbia alone. 98% of them died indoors, and more than half lived alone. The warnings went out, but what didn't happen fast enough was someone checking on them. Neighbourhood teams keep lists of the people most at risk, but a phone tree run by one volunteer on a hot afternoon doesn't scale.

What it does

Doorstep is a Good Neighbor agent built with the Strands Agents SDK. It watches National Weather Service alerts for a volunteer group's area. Each hazard is a plug-in profile: extreme heat is fully tested, whereas extreme cold, smoke and power shutoffs are roadmap profiles. When a warning hits, it:

  • ranks the group's opt-in list by risk (age, living alone, no AC)
  • phones every resident with a natural voice agent (Amazon Nova 2 Sonic), in English or Spanish
  • classifies each check-in, sends cooling-centre info, and assigns volunteers for water or visits

It interrupts the block captain only for the decisions a human must make: an urgent red flag, a high-risk neighbour who won't answer, or a need it can't meet. We replayed the real June 2021 NWS Portland Excessive Heat Warning against a fictional list of 48 residents. The first call started 7 seconds after the alert. All 48 were reached or escalated, which projects to 27.9 minutes on 6 phone lines, with 18 captain decisions against 204 automated actions. The same list is roughly 3.2 hours of phone calls for one volunteer.

How we built it

  • Strands Agents SDK (Python):
    • a Graph (assess → triage → outreach)
    • structured-output agents for classification
    • a dispatcher with tools
    • Strands interrupts with session persistence, so the agent pauses for the captain and resumes after a Telegram tap, even in a different process
    • Cedar authorization on every tool call (allowlisted calls only, minimal disclosure, no agent-initiated emergency calls)
    • hooks for a full audit trail
    • a BidiAgent for voice
  • Amazon Bedrock AgentCore: Runtime hosts the coordinator and the browser voice agent; Observability traces every run. Resident preferences ("hard of hearing — speak slowly") come from the roster in DynamoDB. AgentCore Memory is on the roadmap.
  • Models: Nova 2 Lite (reasoning), Nova 2 Sonic (voice), Nova Micro (simulated residents for drills and evals).
  • AWS: Lambda, API Gateway, SQS, DynamoDB, EventBridge Scheduler, S3, CloudFront, SSM, CloudWatch, CDK. (The phone bridge runs on the operator's machine behind ngrok; it is not hosted.)
  • Channels: Twilio Media Streams (phone), Telegram (captain and volunteers), React dashboard.
  • Evidence: Strands Evals with 40 simulated residents, including hidden red flags and prompt-injection attempts. Red-flag recall 90% over 30 urgent check-ins (every miss was a call where the simulated resident never said the red flag), 20 of 20 forbidden actions denied by policy, 0 policy violations.

Challenges we ran into

  • Getting the page out during the call. A deterministic red-flag backstop listens to the live transcript. On a real phone call, the captain's Telegram message went out 15 seconds before the resident hung up
  • Resuming a paused agent in a different process. The captain might tap twenty minutes later, after the process that asked has died, and exactly one volunteer task has to go out: not zero, not two
  • Understated red flags. A resident who says "fine, bit foggy, I put the milk in the oven" is confused, and the model read it as OK. Our evals caught it; we changed the prompts and added phrases to a deterministic backstop that can only raise a classification, never lower it
  • Some simulated residents never said their red flag at all, and no classifier can catch that

Accomplishments that we're proud of

  • A real phone call that pages a human mid-conversation
  • A replay of a real 2021 warning: 48 of 48 residents reached or escalated, 18 human decisions
  • 0 policy violations, and 20 of 20 red-team attempts refused with an audit reason

What we learned

A single eval run lies, so we ran every urgent persona three times. The model should propose, and deterministic code should decide: state changes, disclosure and permissions live outside the model.

What's next for Doorstep

  • Pilots with a senior building and a neighbourhood emergency team
  • More hazard profiles: extreme cold, smoke, and power shutoffs — each is a profile file, not a rebuild
  • Opt-in enrolment by phone

Built With

Share this project:

Updates

Submission history