Inspiration
What it does
How we built it
Challenges we ran into
Accomplishments that we're proud of
What we learned
What's next for Good Neighbor Recovery Network
Inspiration
Surplus food recovery is not mainly a matching problem. It is an execution problem under changing constraints: a recipient loses capacity, a cold-chain vehicle becomes unavailable, a volunteer route slips, or a recovery plan crosses an approved budget. Communities need more than a chatbot that describes the problem; they need a system that safely changes the plan and leaves an auditable record.
Good Neighbor Recovery Network asks a narrow, important question: can a multi-agent system recover a time-critical food allocation after a real-world failure without inventing capacity, bypassing policy, or silently spending more money?
What it does
The prototype coordinates an 80-meal synthetic surplus-food incident across four community organizations using three distinct Strands agents:
- AllocationAgent proposes the initial multi-organization allocation.
- LogisticsAgent records the versioned capacity failure.
- RecoveryAgent preserves valid work and proposes the smallest safe patch.
The initial plan allocates A30/B35/C15. When organization B drops from 35 to 15 meals of capacity, 20 meals become overcommitted. The system preserves 60 valid meals and proposes A30/B15/C25/D10.
The simulated route costs USD 32 while the autonomous limit is USD 20, so the agent cannot execute the patch by itself. It stops at a persisted human decision. Approval resumes through a separate graph invocation and deterministic policy validation before the plan is committed.
The judge interface exposes Operations, Activity, and Decisions views so the outcome, tool trace, state versions, and human boundary can be understood without reading source code.
How we built it
The agent graph uses Strands Agents SDK 1.52. Each agent can call only a narrow tool for its responsibility; agent text cannot mutate operational state directly. A deterministic policy layer checks inventory conservation, organization capacity, demand, lot-specific cold-chain requirements, delivery-time margins, state versions, idempotency keys, and the autonomous budget.
A compare-and-swap reference ledger makes commits atomic and preserves prior results for duplicate operation keys.
The application is deployed to Amazon Bedrock AgentCore Runtime in Sydney as a Python 3.12 Direct Code package. Runtime version 1 and its DEFAULT endpoint are READY, MMDSv2 is required, and AWS IAM authorizes invocation.
We preserved two real cloud invocations:
- Without approval, AgentCore returns HTTP 200 but stops at pending_decision, executes no recovery commit, and reports zero policy violations.
- With approval, AgentCore returns HTTP 200, runs the separate recovery-resume graph, commits the 80-meal synthetic plan, and reports zero policy violations.
CloudWatch independently records both invocations as successful. The committed evidence omits account and request identifiers while preserving response hashes, business responses, and tool traces.
Challenges we ran into
The hardest problem was separating plausible agent output from authorized action. We moved hard constraints into deterministic tools and policy hooks, required idempotency keys for every mutation, and ended the run when human approval was missing.
AWS account initialization also exposed multiple zero-valued AgentCore quotas. After agent-count capacity became available, the Docker image-size quota was still zero. We used AgentCore Direct Code deployment instead: a Linux ARM64-compatible 30.3 MB ZIP from the exact tested dependencies, uploaded to a private Sydney S3 prefix with the Runtime role limited to read-only access for that prefix.
Accomplishments that we are proud of
- A real, bounded three-agent Strands graph rather than decorative agent roles.
- A fail-closed human budget boundary demonstrated locally and on AgentCore.
- Two HTTP 200 cloud invocations with complete traces and CloudWatch evidence.
- 30 passing automated tests and five consecutive clean graph runs.
- A 60-case deterministic benchmark with committed seeds and raw per-case results: 100% policy-safe outcomes, 100% full recovery on mathematically feasible cases, 100% duplicate-event protection, zero policy violations, and 980 more safely planned meals than the static greedy baseline.
- Honest impact language: 80 meals are secured by a synthetic, policy-valid plan; verified physical deliveries remain zero.
What we learned
Agent safety is most convincing when it is visible as system behavior. The important moment is not the successful final allocation; it is the earlier moment when the agent refuses to execute a feasible plan because it lacks spending authority. Recovery quality should be measured only on feasible cases, while infeasible cases must remain safe and explicit rather than counted as fictional successes.
What's next
The next production step is to replace the in-memory reference ledger with DynamoDB conditional transactions while preserving the same version and idempotency contract. Real deployments would add authenticated organization onboarding, verified pickup and delivery events, notification integrations, route-provider pricing, and measured usability studies with coordinators. None of those production capabilities or physical outcomes are claimed by this prototype.
Reproduce it
Clone the public repository, create a Python environment, install the development dependencies, and run python -m pytest -q. Run python scripts/run_benchmark.py for the 60-case report, then python -m good_neighbor.demo_server --open for the judge interface. No AWS credentials are required for local judging; the private AgentCore endpoint is intentionally not exposed as a public unauthenticated service.
Log in or sign up for Devpost to join the conversation.