Inspiration
Small nonprofits run on volunteers, and volunteers are always in the middle of something else.
Picture a food bank on distribution day. Someone is setting up chairs when the fortieth message of the morning arrives asking where to park. They answer it, because somebody has to. Meanwhile a donor who gave every month for two years quietly stopped in March, and nobody has opened the donor list since.
Two jobs. Both relationship work. Both dropped. Not through carelessness, but because the people who could do them are busy running the thing everyone showed up for.
The obvious fix is to hand it to an AI agent. That runs straight into a second problem, and AWS named it before I did. Their Public Sector guidance on nonprofit agentic AI says the blockers are high volunteer and staff turnover, board accountability, and agents whose decisions cannot be explained.
That reframed the whole project for me. The hard part was never making an agent capable. It was making one a nonprofit board would actually switch on.
What it does
Foyer runs the front desk for a nonprofit that cannot staff one. It does two jobs.
Front Desk answers the questions that arrive over and over. Event dates, times, location, parking, how many spots are left. It also writes RSVPs straight to the roster and refuses politely when an event is full. These actions are small and reversible, so it does them alone and never interrupts anyone.
Donor Steward reads giving history and spots donors who have drifted past their own giving rhythm. Then it drafts a re engagement message in the organisation's voice. And then it stops.
It stops because it cannot do anything else. There is no send tool in its registry. The draft goes into an approval queue with a plain language reason attached, and waits for a person. When a human approves it, the message is emailed and that decision joins the same log as the agent's own actions.
The result is an agent that runs quietly and only surfaces when there is a real decision to make.
Live at usefoyer.online, with the approval panel at usefoyer.online/admin.
How we built it

Foyer is built on the Strands Agents SDK using the agents as tools pattern. An orchestrator reads each inbound message and delegates to one of two specialists. It never answers anything itself, so routing is a decision it makes and logs rather than a keyword match.
The two sub agents have deliberately different tool registries:
- Front Desk holds
get_event_infoandmake_rsvp - Donor Steward holds
get_donor_profile,check_lapsing_donors,draft_followup_messageandqueue_approval
There is no send tool anywhere in that second list. That absence is the whole design. A prompt instruction can be argued with. A missing tool cannot.
The same principle applies on the human side. Approve, reject and send live in
governance/ as plain Python functions, not @tool decorated ones. Only the web
layer can call them, so the agent cannot approve its own draft and cannot send it.
Where the data comes from. This matters, because an agent that invents its answers is no use to anyone. Foyer reads three JSON files. Events hold date, time, location, parking, capacity and spots taken. Donors hold total given, gift count, last gift date and average days between gifts. RSVPs are written when someone signs up.
The model never sees those files. A tool opens the record, does the arithmetic in Python, and hands back a short string. The model writes prose around it. So when Foyer says a donor last gave 194 days ago against a 45 day pattern, that number was computed, not generated.
The clearest proof is a donor it does not flag. One donor in the sample set has given only once, so her average gap is null and no threshold applies. There is no rhythm to break yet. Lapse detection is per donor reasoning, not a fixed cut off someone picked.
Riverside Community Trust is a sample organisation. The donors and events are sample records. The behaviour is real, the nonprofit is not.
The rest of the stack: FastAPI for the server, Groq for inference through the OpenAI compatible provider, Resend for delivery on approval, Railway for the backend and Vercel for the frontend.
Challenges we ran into
Bedrock would not run on my account. Every model invocation returned
ValidationException: Operation not allowed. Not one model, all of them,
including Amazon's own Nova, across every US region. The control plane worked
fine, so ListFoundationModels returned the full catalogue while Converse
refused. IAM had full Bedrock access. The model access page has been retired, so
there was nothing left to request.
It turned out to be an account level activation flag that only AWS Support can clear. A case is open. I lost most of a day to this before accepting it was not mine to fix.
AgentCore got further, then hit a quota of zero. Since AgentCore Runtime is
model agnostic, I reasoned Foyer could run on it while calling a non Bedrock
model. That was right. I worked through four permission walls, the execution role
was created, dependencies cross compiled for ARM64, a 47MB package built and
uploaded to S3, and then CreateAgentRuntime failed with maxAgents limit
exceeded on an account with zero agents deployed.
So the repository is AgentCore ready, with the entry point and configuration committed, but it is not deployed. I would rather say that plainly than imply otherwise.
Keeping the boundary honest under a weaker model. Running on an open weights model made tool chaining less reliable than it would be on a frontier model. The useful discovery was that it did not matter, because the boundary is structural. The agent could not send even when it got confused, because there was nothing to call. That is exactly the property I wanted to test.
A double click sent two emails. decide() did not check whether an item was
still pending, so clicking Approve twice approved twice and sent twice. I caught
it reading a decision log that had the same entry recorded four times. Fixed with
a server side status check and a client side in flight guard. The log is what
caught it, which felt like the project proving its own point.
Accomplishments that we're proud of
The boundary is structural, and anyone can verify it. Open
agents/donor_steward.py and count the tools. Open governance/queue.py and see
that approve is not decorated. Both checks take under a minute. That is a much
stronger claim than "we take safety seriously."
Per donor reasoning, not a threshold. Measuring each donor against their own giving rhythm means a monthly giver is late at 45 days and a twice yearly giver is not. The first time giver who is never flagged is the detail I am happiest with.
One trail, readable cold. Agent actions, human decisions and delivery results all land in the same log in plain language. A board member who has never touched the system can read it without a briefing.
It genuinely runs end to end. A question comes in, an answer goes out, a spot comes off a roster, a draft is written, a human approves, an email arrives, and every step is recorded. Deployed, public, and open source under MIT.
What we learned
Restraint is a feature, and it is the hard one to build. Making an agent do more is easy. Making it reliably stop, in a way that survives a confused model and a persuasive prompt, took more thought than everything else combined.
Put the limit in the architecture, not the prompt. Every guardrail I have written before was a sentence asking a model to behave. Removing the capability entirely is a different category of guarantee. Nothing to reach for beats being told not to reach.
Write logs for people, not for parsers. I started with structured fields
because that is the habit. The log only became useful when it read like sentences.
"James Okafor last gave 194 days ago against a 45 day pattern" does work that
{"donor_id": "d2", "days": 194} does not.
Model agnostic is not a marketing line. When Bedrock closed, switching providers touched exactly one file. The agents, the tools, the boundaries and the logs were untouched. I would not have believed that until I had to do it at speed.
Say what you did not do. The AgentCore situation was tempting to describe vaguely. Writing it out plainly turned out to be easier to defend and, I suspect, more useful to anyone reading.
What's next for Foyer
Meet people where they already are. SMS and WhatsApp rather than a web chat. AWS's own funded nonprofit projects reach underserved communities through low tech channels, and a form on a website is not one.
Real records. Durable storage and a CRM connector to replace the JSON files.
Every tool interface stays the same, since only store.py changes.
A pilot with one food bank. Sample data proves the behaviour. Only a real organisation proves the value, and that means sitting with the person who currently answers those messages by hand.
Finish the AgentCore deployment once the account quota clears.
The brief asked for an agent that runs quietly and only surfaces when there is a real decision to make. That is not a constraint to engineer away later. It is the thing that makes an agent adoptable by the people who have to answer for it.
Built With
- amazon-bedrock-agentcore
- css
- fastapi
- groq
- html
- javascript
- python
- railway
- resend
- strands-agents
- uvicorn
- vercel

Log in or sign up for Devpost to join the conversation.