Inspiration

Some time ago in India, a flood due to mining occurred in Assam. Many Small NGO's tried to help but they can't because of lack the system to approach this scale of disaster coverage which made them very late in helping the affected. Also many of the food items were expired and damaged mainly because of lack of inventory management. Hence, using AI to help these kind of organizations will also help the people who are affected and need help. This was the inspiration for this project and a emotional one.

What it does

OpenShelf is an autonomous background agent for small community organizations — food banks, shelters, mutual aid networks — built for the Good Neighbor Agents track. Instead of another app a coordinator has to open and check, it runs on a schedule and handles three things quietly:

  • Volunteer coverage — checks upcoming shifts against confirmed headcount, recruits from a backup pool, and auto-confirms fills on its own.
  • Inventory & expiry — flags low stock, and routes perishable items nearing expiry to a partner org before they spoil.
  • Donation intake — matches new donation offers to the right partner org based on category, capacity, and need.

It only interrupts a human when a decision genuinely needs judgment — a safety concern, a resource conflict, a policy threshold breach, or real ambiguity between two good options. Everything else gets logged and rolled into a short weekly digest. Every escalation is a single, specific decision (a title, context, and 2-3 numbered options) formatted the way a text message reply would work — because the whole design bet is that the agent is judged by how rarely it needs you, and how good the ask is when it does.

How we built it

The agent is a single strands.Agent with 13 purpose-built tools over a real SQLite datastore (volunteers, shifts, inventory, partner orgs, donation offers, escalations, activity log) — not a thin prompt wrapper. The autonomy policy (what the agent may decide alone vs. must escalate, with explicit dollar thresholds and category rules) lives in the system prompt, and escalate_to_human is the single tool that surfaces a decision to a person.

We built a deterministic, non-LLM implementation of the exact same policy alongside the real agent — this let us prove every piece of tool logic and database interaction works correctly independent of any model, and it doubles as a safety-net fallback in the dashboard whenever Bedrock credentials aren't configured.

The stack: Strands Agents SDK for the agent and tool layer, Amazon Bedrock for the model (with support for both the standard Converse API and Bedrock-Mantle for Gemma 4), Flask for the dashboard API, and a from-scratch matte-themed frontend (Fraunces + IBM Plex Sans, muted olive/clay/mustard palette) instead of a generic SaaS dashboard template. Deployment targets Amazon Bedrock AgentCore Runtime, with EventBridge Scheduler driving the three cycles on a real cadence instead of anyone opening the app.

Challenges I ran into

  • Model access kept moving. I iterated through Claude, Llama 4, and finally Gemma 4 depending on what was actually available in our AWS account — and discovered partway through that Gemma 4 isn't reachable through the standard Bedrock Converse API at all. It's served through a separate OpenAI-compatible endpoint called Bedrock-Mantle, which needed a different code path (Strands' OpenAIModel + BedrockMantleConfig) and an extra dependency we only found by actually running the code and reading the traceback.

  • A real duplicate-escalation bug. Running the same cycle multiple times before a coordinator resolved anything created a new escalation every time for the same unresolved shift gap — exactly the alert fatigue this project is supposed to prevent. We fixed it by adding a stable reference_key to every escalation and checking for an existing pending one before creating a duplicate, with a test that explicitly proves it.

  • Verifying "is this actually the agent talking, or just the fallback?" — we built an honest badge into the dashboard ("Live LLM agent" vs. "Deterministic policy path") that surfaces the real exception name on failure, so it's never ambiguous which engine actually ran.

Accomplishments that we're proud of

  • A genuinely non-trivial agentic implementation — the model reasons through real tool calls (check coverage → find backups → notify → assign, or escalate) rather than following a fixed script, and we verified this by reading actual tool-call traces, not just trusting the summary text.

  • A policy that's honest about its own limits: exact dollar thresholds, exact unit thresholds, and a safety-category list that always escalates regardless of how clean a match looks.

  • A dashboard that's designed around the product's actual thesis — the "needs your call" queue is the dominant visual element because that's the only thing this system should ever ask a coordinator to look at.

  • Catching and fixing the duplicate-escalation bug ourselves, with a test, before it became a demo-day surprise.

What we learned

  • Building "an agent that knows when not to act" is a harder and more interesting design problem than building one that acts — most of the real engineering work went into the escalation policy and dedup logic, not the happy path.

  • Model availability and API surfaces on Bedrock are still moving fast enough that assumptions from even a few weeks earlier (manual model access toggles, which models share which API) can be wrong — worth verifying directly against the current docs and your own account rather than trusting memory.

  • Smaller/newer models are meaningfully less predictable at strict multi-step tool orchestration than larger ones, which reinforced why the deterministic fallback and explicit reference-key dedup weren't optional extras — they're what makes the system trustworthy regardless of which model is behind it that day.

  • I also learned that making a project for a cause is so much more motivating then making a product.

What's next for OpenShelf

  • For the next step I want it to scale to handling more works other then the volunteer shift coverage, inventory & expiry and donation. So more autonomous tasks that are time taking and ultimately fall back to manual checking and verifying which are inefficinet for a community organisation.

-Also, this is as of now is limited to a food bank or a NGO service and cam be modified to work for schools, medical orgs etc.

Built With

Share this project:

Updates

Submission history