Inspiration

One missing volunteer can unsettle an entire community kitchen service. The qualified replacement may already be packing meals, and moving the driver could leave neighborhood pickup uncovered. A coordinator needs a workable update, an explanation, and files that can actually be used.

We chose this specific problem for Good Neighbor Agents: helping small community organizations handle a disruption without quietly shifting risk onto volunteers. CoverCrew was built for this hackathon with AI-assisted development and standard open-source libraries. Riverside Community Kitchen, its volunteers, and every example are synthetic. We have not conducted a field pilot or measured savings for a real organization.

What it does

CoverCrew turns an absence note into a verified replacement roster or a focused request for human judgment. The workbench puts the original assignments, actual tool activity, validation checks, and downloadable results together.

In the main example, Maya cannot cover kitchen lead. Priya has the required food-safety skill but is assigned to packing. The agent moves Priya to kitchen lead and assigns consenting, available Sam to packing. Jordan keeps the locked neighborhood pickup; the welcome desk and cleanup remain unchanged. Two assignment changes restore all five roles.

An ambiguous “Patel” cannot authorize a change when two volunteers share that surname. When no qualified person is available for the whole shift, the agent records the unresolved issue and a next action. Both cases leave the roster unchanged and produce an exception record for the coordinator.

How we built it

The Python Strands Agents SDK, an open-source project from AWS, runs the actual model-and-tool loop. A Strands Agent has five typed tools: inspect_roster, interpret_absence, validate_repair, commit_repair, and record_exception. Tool results return to the model so it can choose its next action. The product does not replay a fixed sequence of fixture outcomes.

Absence interpretation requires known volunteer IDs and an exact quote from the input. Deterministic Python checks enforce complete coverage, skills, availability, absence, consent for changed assignments, overlapping work, hour caps, and locked roles. Individual eligibility summaries help the model find candidates while leaving complete-plan validation mandatory.

Committing invokes validation again and writes roster.csv, calendar.ics, roster.json, audit.json, and packet.zip. File sizes and SHA-256 hashes are checked before the application reports completion. Repeating an identical commit does not rewrite the files.

FastAPI connects this same asynchronous execution path to a JavaScript, HTML, and CSS workbench. Credentials stay on the server. The model route is configurable; our verified configuration uses Strands OpenAIModel with GLM-5.3-Flash and its documented low reasoning setting. This build does not claim an AWS cloud deployment.

Challenges

Language can hide uncertainty: a surname may identify two people, a cancellation may be retracted, or “find cover if one exists” may qualify the repair rather than the absence. We added tests for those distinctions and conservative handoffs when the evidence is insufficient.

Model service compatibility and latency also mattered. We verified the real provider interface, bounded model and tool calls, and treated timeouts as failed runs. Clear tool descriptions and eligibility facts reduced unnecessary calls without bypassing the model's decisions or the final validator.

Accomplishments

Three final live synthetic runs produced the expected outcomes: a two-person repair, a no-safe-cover exception, and an ambiguous-name exception. In that single local verification pass, they took 40.80, 33.40, and 21.77 seconds respectively, with four, three, and two actual tool executions. These are observations from three runs, not a general performance benchmark.

We also verified artifact contents and hashes, repeat-commit behavior, missing-file detection, unsafe proposals, and sanitized error handling. The complete suite contains 101 passing tests plus eight subtests across the scheduling rules, agent boundaries, and HTTP lifecycle; static lint and type checks passed.

What we learned

A useful agent needs a clear boundary between facts, decisions, and permission to persist. Small, inspectable tools make that boundary visible. Giving the model validated candidate facts helps it reach a useful outcome, while the coordinator can still inspect exactly what changed and why a run stopped.

What's next

We want to evaluate the workflow with community-kitchen coordinators before making claims about real operational value. Priorities include roster imports, clearer maintenance of volunteer consent and availability, broader tests of absence wording, and measuring whether the exported packet fits existing routines. A pilot would establish the organization's approval policy before using real schedules.

Built With

Share this project:

Updates

Submission history