Inspiration

The station is not empty. It is empty at 2pm on a Tuesday.

Two thirds of America's firefighters are volunteers, and most rural ambulances are staffed by volunteers too. They have day jobs, often 30 miles from the station. The roster looks fine on paper and is empty in the middle of the workday, and the chief finds out when the tone drops and nobody answers.

One paper changed what I built. Predictive Dispatch of Volunteer First Responders (2023) measured that volunteers answer between 17 and 47 percent of alerts, and that a model trained on a department's own history predicted who would actually respond at 79 percent accuracy.

That is the whole idea. Availability is not turnout. A member saying they are free is not a promise to answer, so four available people can be a legal crew on paper and no crew at all at two on a Tuesday. Turnout scores the probability, not the headcount.

The rest of the picture:

Every other tool tells the chief who is coming. Turnout makes sure someone is.

What it does

1. Roll call by text. Each morning, one message per volunteer: "Around Thu 8a to 5p? Y or N." No app to install, because the people who answer calls at 2am will not install one. It reads "till 2", "morning only" and "sorry, can't" correctly, and honours STOP immediately.

2. It sees the hole before it hurts. A deterministic engine scores every window that cannot make a crew, using the department's own 12 months of call history, per member response probabilities, and live National Weather Service alerts.

3. It closes what it can itself. It asks only the members most likely to say yes for that specific window, inside their quiet hours and under a weekly ask limit.

4. Then it asks the next town. Each department runs its own agent. Millbrook's agent asks Riverton's and Cedar Hollow's agents directly over the Agent to Agent protocol, and ranks the offers by delay, mutual aid balance, and the neighbour's own risk.

5. One text to the chief. Windows that share an answer are batched into a single message. Nothing is confirmed with a neighbour until she replies. It interrupts her once, and the trace shows the times it decided not to.

6. After the call, the report writes itself. A voice debrief becomes a draft NERIS report, the federal format that replaced NFIRS in February 2026, with the uncertain fields flagged rather than guessed.

And it is answerable to the volunteers. There is a page showing, for every member, how many times the agent has asked them this week against the cap, the hours it will not text them, every message it sent including the ones it held until quiet hours ended, and the response history behind the two people it chose. An agent that asks unpaid people for hours of their life should have to show its work to them, not only to the chief.

How we built it

Two boundaries, two protocols.

Inside a department the work is a Strands Graph with conditional edges: Watch, then Closer, then Neighbor, then Chief Gate. A Swarm would have let the agents hand off freely, which is wrong here. When a department asks why nobody was covering Tuesday, the answer has to be the same every time, so each edge reads the shared store rather than the previous node's prose.

Between departments it is A2A, because departments are genuinely separate organisations. Each publishes an AgentCard and answers coverage questions about itself and nothing else. Ask a peer for its roster and no roster comes back.

Policy is enforced in code, not in a prompt. Quiet hours and the weekly ask limit run as Strands hooks on BeforeToolCallEvent. A rule that a prompt can talk its way around is not a rule, and this one is about not bothering unpaid volunteers at 11pm.

The risk score is not a model output. It is plain Python. The model decides what to do about the number; it does not produce it.

On AWS: Amazon Bedrock (Claude Sonnet 4.6 for reasoning, Haiku 4.5 for parsing), Bedrock AgentCore Code Interpreter, where every risk score is computed, and AgentCore Memory, one per department, holding each member's response history. Both are live on the deployment, and each gap card on screen says which path computed it. The service runs on AWS App Runner as one long running container, because the demo holds shared state and two judges pressing the same step on Lambda would land on different instances holding different days.

It is live, with no login. Press the steps in order and the week plays.

Challenges we ran into

Availability is not turnout. The first version counted heads and was confidently wrong. Rebuilding around response probability is what made the product real, and it came from the literature rather than from me.

A2A across an organisational boundary needs identity. Riverton commits an apparatus and a place in the mutual aid ledger on the strength of a request, so it has to be able to tell that Millbrook actually sent it. Every request and confirmation now carries an HMAC over its own contents, checked before the receiving agent evaluates anything. A forged request is refused with a reason. Departments already sign a paper mutual aid agreement naming both parties, so a shared key mirrors how this works on the ground.

Escalation is asynchronous, so Interrupt did not fit. Strands has an interrupt surface that pauses a run and resumes it with an answer. This agent escalates to a chief who may be asleep. The run ends, and her reply is applied hours later. Choosing not to use the obvious feature was the right call, and I wrote down why.

The number was worse before I measured it. The reply parser read 92.6 percent of a corpus I had not written, because its heuristics matched substrings anywhere in a message: "I can't even move" read as yes, because it contains "i can". A length guard and whole word matching took it to 98. I would never have found that testing against my own examples.

Accomplishments that we're proud of

  • Two departments' agents negotiate mutual aid with each other, over a real protocol, in separate processes, and a peer refuses a forged request rather than answering it.
  • It interrupts the chief once. Batching windows that share an answer, and showing the decisions not to interrupt, is the entire point of a background agent.
  • Measured rather than asserted: it refuses 14 of 14 adversarial cases and wrongly refuses 0 of 6 legitimate ones. That second number is the one that matters, since refusing everything would score perfectly on the first.
  • It is accountable to the people it asks, not just to the person it reports to.
  • Every page works at 200 and 400 percent zoom, in both themes, with a keyboard.

What we learned

Put the policy in the code. Anything that protects a person from the agent belongs somewhere a prompt cannot reach.

Pick the agent shape from the problem, not the menu. A conditional Graph inside an organisation and A2A between them are different answers because they are different questions.

Say what is designed and what is running. AgentCore Code Interpreter and Memory are live here. Runtime, Gateway and Identity are designed and labelled as designed. Being precise about that is more persuasive than claiming all six.

What's next for Turnout

  • Real messages through AWS End User Messaging. The carrier is simulated so the demo plays identically every time, but the same code path, policy hooks and segment checks sit above it.
  • AgentCore Runtime, Gateway and Identity, moving each department into its own isolated runtime with scoped credentials.
  • NERIS submission with credentials issued to a named officer. It drafts and then stops today, which is deliberate: a person should read what is about to enter a federal incident record.
  • A pilot with one real department, and more state ratio and mutual aid rules.

Everything in the demo is synthetic. Millbrook, Riverton and Cedar Hollow are fictional, and the members, call history and incident are generated. The weather tool calls the real National Weather Service API.

Built With

  • a2a
  • agentcore-code-interpreter
  • agentcore-memory
  • amazon-bedrock
  • amazon-bedrock-agentcore
  • amazon-ecr
  • aws-app-runner
  • claude
  • docker
  • fastapi
  • javascript
  • national-weather-service-api
  • opentelemetry
  • pydantic
  • python
  • strands-agents
Share this project:

Updates

Submission history