Inspiration

Blind and low-vision people run into small, situational tasks all the time that need a sighted person for a moment: reading a label, describing a document, figuring out which side of the street a bus stop is on. Volunteer networks already exist for this, run by a library, a nonprofit, an informal neighborhood group. But the coordination is almost always manual. Someone has to notice the request, remember who's free, follow up, and know when to step in themselves instead of waiting. That gap in coordination, not the underlying task, is what actually breaks down. We wanted to build the dispatcher, not another assistant.

What it does

A requester submits a task in plain text plus an optional photo. The agent:

  1. Answers directly if it can do so confidently from what's given, no volunteer gets pulled in for something the agent already knows.
  2. Otherwise finds an available volunteer with the right skill and notifies them by email.
  3. Escalates to a human coordinator if no volunteer is available, or if the request reads as safety-critical (medication, an unsafe location), even when a volunteer is free.
  4. Logs the outcome either way.

It runs autonomously end to end. A human only sees it when there's a real decision to make.

How we built it

The agent runs on the Strands Agents SDK in Python: one Agent with a system prompt that encodes the decide/route/escalate judgment instead of hardcoded branching. For the model we went with thegrid.ai's agent-prime instrument through Strands' OpenAI-compatible provider, mainly to avoid needing a billed AWS account for a hackathon side project. The tools are find_volunteer, notify_volunteer, escalate, and log_outcome, all backed by real sends through Resend. State lives in a JSON-backed volunteer roster and a pending.json file that tracks requests waiting on a response.

For autonomous escalation, a GitHub Actions workflow runs on a 5-minute cron and calls a standalone, deterministic script (escalation_check.py, no LLM call) that sweeps pending.json for anything overdue and escalates it on its own. GitHub-hosted runners are ephemeral, so the job commits its own state changes (pending.json, outcomes.log) back to the repo, which is how the next run picks up where the last one left off.

Challenges we ran into

The original spec called for Amazon SES for volunteer notification, carried over from an early Bedrock-based sketch. We caught this before writing any code against it: SES needs the same billed AWS account we were deliberately avoiding, so we swapped to Resend instead.

Resend's sandbox tier only delivers to the account owner's own registered email until a domain is verified. Rather than put a real personal address in the public seed data (volunteers.json), we had the roster override the matched volunteer's email with a gitignored DEMO_INBOX environment variable at runtime. The committed data stays generic, but local runs and the demo still deliver real email.

A subtler one: two environment variables were read as module-level constants, so they got evaluated before load_dotenv() ran and silently fell back to placeholder values. It only surfaced because we tested actual email delivery end to end instead of trusting the code on inspection. Moving the reads inside function bodies fixed it.

Our own .gitignore's blanket *.log rule almost silently excluded outcomes.log, a file the escalation workflow depends on being tracked since it's how state survives across ephemeral CI runs. We caught it with a dry-run git add -n . before the first push, not by assuming it would work.

Recording the demo video hit its own wall: this machine runs macOS 12.6, and the render pipeline's managed headless-Chrome dependency only ships builds for macOS 13 and up. We pinned an older Chrome-for-Testing build that still supports this OS instead.

Accomplishments that we're proud of

All four acceptance criteria are proven with real runs, not mocked: real thegrid.ai model calls, real Resend emails landing in a real inbox, and a real live GitHub Actions run of the escalation cron. That includes the safety-critical case, where the agent escalated a medication request without even attempting to match a volunteer, despite one being available on the roster. The demo video is built the same way. Every terminal output, log line, and timestamp on screen is copied from an actual run against live services, not staged.

What we learned

The coordination logic, knowing when not to bother a human and when to stop waiting and escalate, turned out to be a more interesting problem than whether an LLM could answer the request. And for a pipeline touching real people's real needs, proving each branch with a falsifiable, real-service test catches failures that "looks right" never would. Half the challenges above only surfaced because we insisted on testing delivery, not just logic.

What's next for Good Neighbor Agent

Phase 2 is a web intake form so requesters don't need direct terminal or API access. Phase 3 adds SMS/WhatsApp notification for volunteers who aren't watching email. Phase 4 moves from a JSON roster to real volunteer accounts and a database. We're doing these in sequence, not in parallel. A small, fully verified thing beats a bigger, half-finished one, especially in a domain where a broken escalation path has real consequences for someone waiting on help.

Built With

  • github-actions
  • json
  • python
  • resend
  • strands-agents-sdk
  • thegrid.ai
Share this project:

Updates

Submission history