Inspiration

Every day, grocers and restaurants have surplus food that's still good, but getting it to a food bank before it expires takes a human coordinator on the phone, juggling who has capacity, who's driving where, and how much time is left. That coordination overhead is the actual bottleneck: 29% of the US food supply went unsold or uneaten in 2024, about 114 billion meals' worth (ReFED), while 47.9 million Americans lived in food-insecure households that same year (USDA ERS). The food and the need both exist, often within a few miles of each other. What doesn't scale is a human being available at the exact moment a donation call comes in, with an up-to-date picture of every food bank's capacity and every driver's schedule.

The Good Neighbor track asked for an agent whose win isn't measured by one person's productivity but by a group benefiting together. A food-rescue router fit that exactly: every successful match benefits three separate parties, the donor, the food bank, and the driver, plus the community the food bank serves.

What it does

Windfall is an autonomous multi-agent system that takes one surplus-food donation offer and gets it to a food bank that needs it, via a volunteer driver who can carry it there in time, with no human approval step in the loop. A Coordinator agent checks the donation against a food-safety handling window, delegates to a Matching Specialist to find the right food bank and a Logistics Specialist to find the right driver, each reasoning independently over live data and real geographic distance, then either commits to a match (and notifies all three parties) or explicitly escalates to a human coordinator with its reasoning, every time.

How I built it

  • Strands Agents SDK, agents-as-tools pattern: the Coordinator never touches food bank or driver data directly, it consults two separate Agents wrapped as @tools. This isn't one agent with two lookup functions, it's real delegation, and it shows in the reasoning: each specialist checks capacity, need level, vehicle capacity, and availability window before answering.
  • Amazon Bedrock (Claude Sonnet) as the model behind every agent.
  • Real geographic reasoning: donors, food banks, and drivers carry real Austin, TX coordinates, and both specialists rank candidates by actual haversine distance, not a coarse zone label.
  • Food-safety constraints: a dedicated tool checks the donation's pickup window against a per-category safe-handling limit, so a geographically perfect match can still get escalated if it would sit too long before pickup.
  • AgentCore Memory for cross-donation recall, so the Coordinator notices when a driver it just committed to another load moments ago comes up again, instead of treating every donation as if it were the network's first.
  • FastAPI + SQLite + Server-Sent Events for a live dashboard: every match, escalation, and activity update pushes to the UI the instant it happens, with a real Leaflet map animating the route.
  • Deployed twice: the identical Coordinator and specialists run in-process behind the public dashboard (AWS EC2, Docker, Caddy for automatic HTTPS) and standalone on AWS Bedrock AgentCore Runtime, invokable independently and verified live with a real agentcore invoke.
  • Dark mode, an accessibility pass (real dialog semantics, live regions, focus management, reduced-motion support), and optional Supabase magic-link sign-in, none of it gating the core demo.

Challenges I ran into

  • A genuine AgentCore Memory bug, not a config typo. Early on, invoking the deployed runtime failed with "Invalid conversation manager state." Root cause: the local dashboard and the standalone runtime shared one fixed memory session ID but used different Strands conversation managers, and AgentCore Memory tries to restore persisted state shaped for the wrong one. Fixed by giving each deployment its own session identity, verified with a real invoke afterward.
  • The same class of bug bit twice. After enough local test donations, the Coordinator's real memory correctly recalled all of them, to the point that even an "easy match" scenario started escalating because every driver looked already committed. Not a bug so much as the feature working exactly as designed, just inconvenient for testing, solved with an overridable session ID for one-off runs.
  • A project path with a space in it broke the AgentCore CLI's CDK asset bundling (it shells out to uv, which mangles the path). Workaround: build and deploy from a space-free copy, then sync the deployment state back.
  • Windows console encoding. The Strands SDK's default callback handler streams raw model tokens to stdout, which crashes on Windows's cp1252 console the moment the model emits certain characters. Needed callback_handler=None on every nested agent, not just the top-level one.
  • Multiple zombie dev-server processes, each running an independent 20-minute auto-reset timer against the same shared SQLite file, silently wiped an in-flight donation mid-recording of the demo video. Diagnosed by walking the actual process tree, not by guessing.

Accomplishments that I'm proud of

  • Everything claimed is actually verified, not just asserted: a real agentcore invoke against the deployed runtime, a real cross-donation memory recall visible in the agent's own reasoning ("Sam R. was already committed to donation..."), a live public dashboard a judge can click into and try themselves rather than a recorded mockup.
  • The multi-agent delegation is real depth, not decoration: the Matching and Logistics Specialists independently check live capacity, distance, and availability, and the Coordinator's food-safety check can override an otherwise-perfect match.
  • Finding and fixing genuine, non-obvious bugs through actual debugging rather than guessing: a memory-session collision between two deployments that only showed up as a cryptic SDK error, and a fleet of orphaned dev-server processes each running its own reset timer against one shared database.
  • Getting the deployment fully production-shaped: HTTPS on a real domain via an automatically-renewing Let's Encrypt cert, not just a bare IP over HTTP.
  • A dark mode and accessibility pass done properly, real dialog semantics, live regions, focus management, reduced-motion support, not a cosmetic toggle.

What I learned

That "agents as tools" earns its complexity when specialists actually reason independently, not when it's a more elaborate way to call the same lookup function. That shared state, a memory session, a database file, a deployment, needs explicit isolation between environments even when it's "the same app," or it produces bugs that look inexplicable until you trace exactly who's touching what. And that verifying against a real headless browser, not just curl, caught UI-affecting bugs that a passing API response would have hidden entirely.

What's next for Windfall

The honest gap right now is real-world validation: the donor, food bank, and driver data is a realistic synthetic stand-in, not a live partnership. The natural next step is reaching out to an actual food-rescue network or food bank coalition to pilot Windfall against real donation data instead of the synthetic Austin network. Beyond that: expanding past a single city, adding SMS or push notifications so donors and drivers don't need to be watching a dashboard, and reusing the same Strands multi-agent coordinator pattern for a second, unrelated domain to prove it's an architecture and not a one-off.

Built With

Share this project:

Updates

Submission history