Inspiration

Every week, at food pantries and shelters everywhere, the same scene plays out: someone cancels a volunteer shift, and the coordinator — usually one overloaded staffer — spends the evening texting twenty people one by one. Roughly 1 in 4 scheduled volunteers no-show, and coordinators describe replying to dozens of volunteers a week by hand, copying answers into a spreadsheet. Volunteer software gave them portals and reminders, but nobody logs in and nobody has time to check — while the actual work, the chase, stayed 100% manual. Paid-staffing tools sell autonomous backfill, but for volunteer-run orgs that product simply doesn't exist. We built the thing the problem has been asking for: an agent that works the phone tree, receipts everything, and surfaces only when judgment is required.

What it does

wpilot watches an org's shift roster. When a shift goes short — a cancellation comes in, or a scan finds unfilled slots — it runs a backfill campaign: it ranks eligible volunteers, texts them one by one with a human-quality offer, waits for replies, books the first YES atomically into the roster, and posts a receipt for every action. The coordinator is interrupted only when a human must decide — the pool is exhausted, or a booking is sensitive (low-reliability volunteer, restricted role). One tap answers; the campaign resumes, even across restarts. Coordinators get one clipboard page (shift cards, a needs-you queue, an activity feed); volunteers never see an account or a login — they answer texts or tap a single-use confirm link. Safety rails (eligibility, adults-only defaults, quiet hours, ask caps, consent verification, atomic writes, undo) are enforced in code, never in the prompt.

How we built it

The deterministic core is Python: a campaign engine over a SQLite roster store, a deny-by-default eligibility policy, an interactive messaging channel, and immutable receipts with pre-booking snapshots. On top sits a Strands Agents SDK loop with 8 tools (gaps, ranking, offers, replies, booking, receipts, read-receipts, undo) and a BookingApprovalHook that raises session-managed interrupts for sensitive bookings, with trust answers persisted in agent state. The same tools serve a FastAPI demo API, a Next.js coordinator console, a live demo console with a simulated volunteer phone, and a no-login confirm page. The live loop runs on Amazon Bedrock via the mantle OpenAI-compatible endpoint (model-portable: Anthropic → OpenAI → mantle → Bedrock default → Ollama fallback). Deployment is all-AWS: the API on App Runner, images in ECR, the UI on Vercel, with a Bedrock AgentCore Runtime target scripted and waiting on account quota. Reliability is proven, not asserted: 34 pytest tests and a 15-scenario eval suite (product guarantees like "minor never texted" and "no booking without a recorded YES"), plus a documented live loop proof — real model, real tools, approval interrupt, process killed, resumed cold, shift filled.

Challenges we ran into

Three hard ones. First, our AWS account sat in verification for the whole build: SigV4 Bedrock inference refused on every model and region, and the AgentCore quota was zero. We proved it was account-level (not auth, not region) and routed the live loop through Bedrock's mantle bearer-token endpoint — after discovering our first key had expired and minting a fresh one. Second, the local small models we used for offline proofs were slippery: one booked a volunteer who said NO, another invented audit receipt kinds, a third recursed a tool into itself. Each became a hardening feature: durable consent verification in the offers/replies tables, an audit-kind allowlist, actionable tool errors, and a regression test per incident. Third, demo honesty: we refused canned scripts, so the simulated SMS channel is a real channel boundary — which meant building timeouts, quiet hours, session persistence, and a reseedable demo org so any judge can trigger a cancellation cold.

Accomplishments that we're proud of

  • A live public demo where anyone can trigger a cancellation and watch a real agent fill the shift — deterministic path and live LLM path.
  • A documented kill-and-resume: approval interrupt raised, process killed, new process resumes from the session store, shift filled.
  • Safety properties most agent demos never attempt: consent-verified booking, deny-by-default eligibility, single-consume undo, full receipts.
  • 15/15 eval scenarios green, each asserting a coordinator-meaningful guarantee, shipped as a report artifact in the repo.
  • All-AWS deployment (App Runner + ECR) with a container-verified AgentCore target ready the moment quota lands.

What we learned

Technically: keep the model proposing and the code disposing — every time we let the LLM own a safety decision it eventually betrayed us, and every hardening made the product better. Session-managed interrupts are the difference between a chatbot and an agent that "runs in the background and surfaces when needed." Product-wise: trust is the feature for this user — receipts, approval gates, and undo aren't compliance overhead, they're the adoption mechanism. And process-wise: evals written as product guarantees ("the minor is never texted") catch real bugs that unit tests miss, twice in this build.

What's next for wpilot

Pilots with five food pantries, gated on hard metrics (median time-to-fill under 30 minutes, zero median coordinator touches per fill). Then: real SMS via a registered provider, Google Sheets as roster source of truth, Cedar policies replacing the rules module 1:1, and per-org configuration. Longer term: portable volunteer reliability memory across orgs, distribution through pantry networks, and $30–100/mo per-org pricing matched to avoided coordinator hours. The wedge is backfill; the moat is the outreach memory that accumulates with every campaign.

Built With

Share this project:

Updates

Submission history