Inspiration
Community fridges fail quietly: food spoils because no one happens to check it, and shelves sit empty because nearby donors don't know they're needed. Today this is "managed" through a muted WhatsApp group nobody reads. We wanted to fix the actual behavior, not add another app volunteers have to remember to open.
What it does
Sentinel is a background agent that watches a single community fridge end-to-end, entirely over SMS — no app, no login. A volunteer texts a photo and short description to the Sentinel number. The agent classifies the state into one of three outcomes:
Risk detected → texts the volunteer to pull the flagged item. Critically empty → finds the nearest donor and sends a specific restock nudge. All fine → logged silently, no one contacted.
That silent path is the point: most check-ins should need zero human involvement, and the system only speaks up when there's a real decision to make.
How we built it
Sentinel runs on the Strands Agents SDK, using the Graph pattern rather than an exploratory Swarm, since the decision path (intake → assess → one of three branches) is fixed and deterministic.
Intake node parses the SMS into structured signals: items, freshness cues, fill level. Assess node is a real Strands Agent with a structured-output model, backed by a lookup_shelf_life tool covering 20+ food categories with shelf life, spoilage signals, and alias matching. Branch nodes text the reporting volunteer (alert_pull) or the nearest donor by real geodistance (alert_empty), via Twilio, through a notifier that fails safely instead of crashing. Donor roster, fridge ID, and credentials pass through invocation_state, never the LLM prompt, so donor numbers never appear in a reasoning trace. Built as a FastAPI service meant to run continuously on AgentCore, not invoked manually.
Challenges we ran into
The hardest problem wasn't wiring pieces together — it was making sure "assess" did genuine reasoning instead of quietly becoming a rules engine matching keywords like "empty" or "bad." We built an evaluation harness scoring the assess node against sample check-ins for accuracy, silent-resolution rate against our >70% target, risk recall, and false-alarm rate, instead of just eyeballing outputs.
We also had to avoid a fallback that depends on the same thing it protects against — a "backup" that still calls the live model doesn't help if the model is what fails on stage. We built a deterministic local fallback the graph uses automatically if the model call errors, so the webhook degrades instead of crashing.
Accomplishments that we're proud of
Our core claim — genuine agentic reasoning, not a rules engine — is something we can demonstrate with numbers, via the eval harness measured against our own product spec. We're also proud of how far the shelf-life tool grew from a five-item dictionary into a real reference table with spoilage signals and alias resolution, and of a graph that never fully falls over when the model provider is unreachable.
What we learned
We learned to design deliberately around invocation_state — what an agent needs to reason over versus what should stay out of its context for privacy and cost. We learned "agentic" has to be actively defended throughout a build, since deterministic shortcuts creep back in under deadline pressure; testing against expected outcomes is the only real defense. And a fallback is only as good as its independence from the failure it protects against.
What's next for Sentinel
Deliberately out of scope for now, but on our radar: fairness-aware donor rotation so the same donor isn't always pinged, a lightweight admin log view, real temperature-sensor ingestion, and WhatsApp Business API support. We're staying away from a volunteer-facing app or multi-fridge orchestration — the fewer things a volunteer has to open, the better it's working.
Log in or sign up for Devpost to join the conversation.