Inspiration

Volunteer coordinators at mid-size nonprofits run 50–500 people with no assistant, and lose 22+ hours a week to the same loop: match people to shifts, chase replies, send reminders, find cover when someone doesn't show, log hours by hand, and produce numbers for the board. None of it needs a human's judgement — it needs a human's persistence. That's what an agent is for. We wanted the coordinator to post a shift on Friday evening and have it staffed, reminded and reported without coming back to the dashboard.

What it does

For the coordinator: post a shift (program, time, place, skills, headcount). That's the last action they take.

For the agents (a Strands multi-agent graph, five roles, 17 tools):

  • Scheduler — reads the shift, calls the matcher (hard filters on required skills and the shift's weekday, then ranks by reliability, past participation and hours), and assigns the top candidates.
  • Communicator — writes each matched volunteer a personal email through Amazon SES containing their exact one-tap /respond link, and logs every send. Also runs the 48-hour and 2-hour reminders.
  • Recovery — fifteen minutes after start, checks who hasn't checked in, contacts ranked replacements, and escalates to the coordinator.
  • Tracker — logs hours from check-in to check-out and adjusts reliability scores so the next match is smarter.
  • Reporter — produces the weekly coverage / no-show / hours report, stored in DynamoDB and S3.

For the volunteer: a public sign-up page, an email with one link, and a phone-friendly confirm page. Or just reply YES / NO — inbound mail flows SES → SNS → Lambda → the API and is applied to the assignment.

For everyone: a background worker fires each action on schedule (schedule → invite → T-48h → T-2h → T+15m no-show check → hours at end), and every tool call — name, input, result, timestamp — lands in an audit table that the dashboard renders live. The coordinator's overview shows only what genuinely needs a decision.

How we built it

  • Agents: Strands Agents SDK 1.52 — five agents, each with its own system prompt and tool subset; hooks on BeforeToolCallEvent (recipient validation, PII pattern blocking) and AfterToolCallEvent (audit write); a reliability hook that re-prompts an agent that stops short of its required tool.
  • Model: Mistral Large 3 (675B) on Amazon Bedrock.
  • Data: DynamoDB (volunteers, shifts, communications, reports, audit) with consistent reads so a volunteer who signed up seconds ago is visible to the very next cycle; S3 for reports.
  • Messaging: Amazon SES for outbound (verified domain, DKIM, SPF/DMARC) and inbound (receipt rule → SNS → Lambda → /api/ingest/email-reply). SMS via SNS is designed and wired but held as a future integration — the account lacks the End-User Messaging subscription, so we made email the first-class channel rather than ship a dead-lettered "sent".
  • Observability: CloudWatch custom metrics (Vshift namespace: shifts coordinated, no-shows detected/recovered, hours logged, communications sent), a dashboard and alarms.
  • Deployment: the agent is packaged and deployed to an Amazon Bedrock AgentCore Runtime; production traffic today is served by the same FastAPI service on a VPS behind Caddy (auto-TLS) with a systemd automation worker. [teammate: confirm current AgentCore quota status before submitting]
  • Frontend: Next.js 15 / React 19 / Tailwind — a public landing page, sign-up and respond pages, and a coordinator app (overview, shifts, volunteers, communications, live audit trail, reports, automation) with a ⌘K command palette. The browser never calls the backend directly; a same-origin route handler proxies to the API at request time.
  • Tests: 61 pytest cases (39 run without AWS; 22 integration tests against live tables).

Challenges we ran into

  • The one that mattered most: our first on-camera take failed. The dashboard sent shift times as …T09:00:00.000Z; the server's Python 3.10 fromisoformat rejects a trailing Z, so the matcher threw and the Scheduler assigned nobody. The audit trail told us exactly where. Fixed in the dashboard (+00:00), and it's now on the backend's list to parse leniently too.
  • The Scheduler stalled on a live test: the model computed the shift's weekday itself, got it wrong, filtered the pool to zero and gave up. Fix: route the prompt through the matcher tool instead of letting the model filter by day, and turn on consistent reads.
  • Email plumbing is most of the work. Domain verification, DKIM, MX for inbound, a receipt rule, SNS, a Lambda bridge, and the SES sandbox — plus discovering that SMS needs a separate AWS subscription and carrier registration.
  • Getting a public HTTPS endpoint on an Oracle VPS meant opening ports at two layers (host iptables and the VNIC security group), then a Caddy certificate that couldn't issue until DNS pointed at the box.
  • Shift lifecycle statuses (in_progress, completed) existed in the model but nothing ever set them, so past shifts stayed "filled" forever. Now the worker advances them.

Accomplishments that we're proud of

  • The demo is not staged. In the film, a real shift is posted, the worker fires 65 seconds later with no manual trigger, the matcher finds exactly 2 of 54 volunteers who speak Spanish, both get SES-delivered invitations with message IDs, one confirms from a phone while the coordinator's screen updates on its own, and the shift reaches 2/2 — eight tool calls, 78 seconds from posting to invitations logged, zero coordinator actions.
  • Every number on screen has a receipt in the audit table, and the dashboard's coverage is counted from confirmations, not from the status field — it refuses to call a shift "filled" until people actually said yes.
  • Guardrails that are hooks, not prompts: nothing leaves without a valid recipient, PII patterns are rejected from tool inputs, and an agent that forgets to call its required tool is re-prompted until it does.

What we learned

  • Give the agent the data, not the lookup. The Communicator has no query tools — its caller pre-fetches addresses and links, so it cannot hallucinate an email address.
  • An audit table is the best debugger you will ever have. Both of our real bugs were found by reading it.
  • Automate the schedule, not the judgement. The worker owns when; the coordinator only sees the shifts that genuinely need a person.

What's next

  • SMS through AWS End-User Messaging once provisioning clears, and WhatsApp for volunteers who prefer it.
  • Custom MAIL FROM domain and SES production access for full deliverability.
  • Auth on the coordinator app (Cognito), keeping /signup and /respond public.
  • Recovery that closes the loop on replacement confirmations automatically, and monthly reports with trends.

Built With

Share this project:

Updates

Submission history