On the Line
On the Line is a community resource directory that phones ahead. It calls the food banks, clinics and shelters it lists, confirms what is still true, publishes only what a provider said out loud and confirmed on readback, and hands everything that needs judgment to a person. When a neighbour reports a closed door, that report becomes the next call in the same database transaction.
Check it while you read
The deployment is live: https://54-84-108-211.sslip.io
- The public side needs no account. https://54-84-108-211.sslip.io/help is the search a person looking for help would use. Take a referral, then report a closed door and watch that report reach the top of the call queue.
- Sign in to the curator console with
judge/ontheline-demo-2026. - On the overview, press Call the provider on the first queued service and listen to the call. Turn the sound on.
- Decisions holds the closure the agent refused to publish, with the evidence diff and the provider's own quote.
- Activity is the append-only chain the database will not let anyone edit or delete.
- Add
?demo=1to any page for a narrated walkthrough that drives the real console. Escape stops it and hands the product back.
The problem, in one trip
Two buses. Two children. A pantry that closed in July and a directory that still says Wednesdays, nine to twelve. The referral was free to send and expensive to follow.
United Way reports that the US 211 network made 19 million referrals across 1.6 million locally curated services in 2025. Open Referral documents the maintenance loop behind that number: fragmented directories repeatedly ask the same service providers for the same updates, staff spend their time re-verifying records, and the people relying on the listing get less reliable information. Keeping those listings true is repetitive work that almost nobody is funded to do, so it does not get done, and the cost lands on the person least able to absorb it.
Who it is for, and why it matters
The same phone call helps three different people, which is what makes this a Good Neighbor project rather than a data-quality tool.
| Who | What they get |
|---|---|
| The person looking for help | A listing a human being confirmed, with the date they said it |
| The volunteer holding the directory together | The routine calls made without them; only judgment reaches their inbox |
| The organization being listed | Called once instead of twelve times, because the answer is published as open HSDS any other directory can consume |
That third row is the part most directory tools leave on the table. Open Referral's own strategic case is that fragmented directories pester the same service staff for the same updates. On the Line exports stable-UUID HSDS 3.2 on a public endpoint, so the result of one call is publishable to every other directory. Call once, publish for everyone.
The mechanic
Someone finds the door locked and says so. In the same database transaction, that service jumps to priority 1000 at the top of the call queue, and the agent phones the provider. The next neighbour finds it open.

Help-seekers are the freshness sensors. The provider calls are the repair crew. A directory chatbot does not close this loop, because it answers the question it was asked and then forgets it happened. Here the failed referral and the new verification task share one transaction, and a direct failure is merged into any existing open task for that service, so the queue never double-counts.
What the call sounds like
The console places the call and you hear it. The agent opens by saying what it is, reads the changed hours back before recording them, and waits to be told it got them right:
On the Line: Let me read that back. Wednesdays, ten in the morning to one in the afternoon. Have I got that right?
Provider: Yes, ten to one on Wednesdays is correct.
That readback is the evidentiary standard, not decoration. Without an explicit
statement and a confirmed readback, the change is held for a person. The audio on
screen is rendered turn by turn from CALL_SCRIPTS, each speaker in a separate
voice through a 300 to 3400 Hz telephone band, and the transcript reveals each
line exactly when it is spoken, so the words you read and the words you hear
cannot drift apart.

How Strands is used
The agents do the interpreting; plain Python decides what gets written.
Voice, both directions: a Strands BidiAgent on BidiNovaSonicModel runs the
provider verification calls and the inbound help-seeker line, with interruption
handling. src/ontheline/audio/telephony_io.py bridges a carrier stream to the
agent: mu-law 8 kHz in from the phone, PCM to the agent, PCM back, mu-law out,
with bounded backpressure so a fast talker cannot outrun the buffer. On a
configured deployment the carrier is Twilio bidirectional Media Streams; every
HTTP callback and WebSocket upgrade is verified against Twilio's signature over
the canonical public URL, and audio is never persisted.
Extraction: a GraphBuilder graph fixes the order, a safety reader then a
structured extractor, entered only after a call has produced a transcript. The
safety reader preserves every provider statement verbatim and flags ambiguity,
distress, a help-seeker answering instead of staff, and stop requests. The
extractor runs on Amazon Nova 2 Lite with structured_output_model=ExtractionResult,
so its output validates against a Pydantic schema before anything downstream sees
it.
A domain contract, not just a JSON shape: the extractor carries a Strands
GoalLoop. JSON validity is table stakes; the goal validator enforces domain
invariants and retries the model until they hold. A non-active status has to set
removes_service=true. A read_back_confirmed change has to also be
explicitly_confirmed. A model that returns well-formed but domain-invalid
output is sent back with specific feedback rather than accepted.
Lifecycle and state: invocation hooks (HookRegistry, BeforeInvocationEvent,
AfterInvocationEvent) record which agent ran and in what order, and the trace is
persisted with the verification. FileSessionManager keys graph state on a
path-safe, task-specific session id, so a retry resumes rather than starting
over.
AgentCore: Dockerfile.agentcore builds a stateless ARM64 runtime exposing
/ping and /invocations. It runs the extraction graph and returns only
schema-valid extraction plus the lifecycle trace. It holds no directory database
and its IAM role grants Nova invocation and telemetry only, so policy, consent
and persistence stay in the application boundary. OTL_AGENT_MODE=agentcore
selects it; bedrock runs the same graph in-process; demo runs a deterministic
offline adapter that escalates anything it cannot establish explicitly.
Evaluation: the official Strands Evals SDK runs a versioned experiment of 12 adversarial provider personas. The committed deterministic report records 100% exact policy matches, a 0% false-apply rate, a 0% over-escalation rate, and successful blocking of every tested removal, missing disclosure and stop request.
The line the model never crosses
Model output is a proposal. A deterministic Python policy in ontheline.policy
classifies complete evidence into one of four tiers before any transaction
touches the directory.
| Tier | Evidence | What happens |
|---|---|---|
| 0 | Every requested field confirmed unchanged | Refresh the assurance date |
| 1 | Allowlisted field, explicit statement, confirmed readback, canonical value | Apply it, with the quote attached |
| 2 | Removal, eligibility, ambiguity, contradiction, low confidence | Hold it for a person |
| 3 | Missing AI disclosure, distress, help-seeker reached, hostility, stop request | End the call, flag, and suppress when requested |
The policy validates field coverage, confidence, explicit confirmation and readback, canonical HSDS schedules and E.164 phone numbers, high-judgment fields, and removals. The model is never handed a database, arbitrary SQL, a publish tool, a delete tool, or consent authority. Changes are atomic per verification: the overall tier is the maximum across every proposed change, so if one field needs judgment, nothing else from that call is applied on its own.
The evidence chain
Every automatic change and every human decision keeps its call, timestamp, quote,
consent reference, policy tier and agent trace in a SHA-256 hash chain. Each hash
covers the prior hash, the canonical payload, the event type, the actor and the
timestamp, and database triggers reject every UPDATE and DELETE on the audit
table. The human-interrupt requirement is a database-backed boundary rather than
a pause inside a model process, which is what lets a pending decision survive a
restart and be resolved by an approve, edit or reject transaction that
revalidates the field before it publishes.
What the build settled
A readback is the input, not a formality. The first version of the call had a two-turn transcript with no readback in it, while the policy required one and the tier table advertised one. Making the call a real conversation, where the agent reads the changed value back and waits to be told it is right, is what lets a hedged or missing confirmation land in the decision inbox instead of the directory. The conversation is the evidence.
A goal is a stronger contract than a schema. Nova 2 Lite returns valid JSON
readily. The failures that matter are domain failures: a closure that does not
mark the service for removal, a readback flag set without an explicit statement
behind it. Encoding those as a Strands GoalLoop goal, with feedback the model
retries against, is what turns a schema-valid answer into a domain-valid one.
The interrupt belongs in the database. Putting the human gate in SQLite, not in the interface, is what makes "only judgment reaches a person" true across a restart, a redeploy and a second worker. An interrupt that lives in a process dies with the process.
A guard that has never been fired once is an assumption. The live site restores its demonstration directory on a timer. Testing that timer once, on purpose, before letting it run unattended during judging, is the only thing that could have caught what it was doing: an early version tore down the stack with its volumes and discarded the Let's Encrypt certificate on every reseed, which would have exhausted the certificate quota and taken the site down mid-judging while every containers-came-back check still passed. It now restores only the data volume and leaves the certificate in place.
Technologies used
AWS and Strands
| Strands Agents SDK | BidiAgent on BidiNovaSonicModel for two-way voice, a GraphBuilder extraction graph, structured_output_model, a GoalLoop domain validator, HookRegistry lifecycle tracing, FileSessionManager resumable sessions, and the Strands Evals SDK |
| Amazon Nova 2 Sonic | live bidirectional speech for the provider and help-seeker conversations |
| Amazon Nova 2 Lite | the structured HSDS extractor behind the graph |
| Amazon Bedrock AgentCore | a stateless ARM64 extraction runtime (/ping, /invocations) that returns schema-valid extraction and trace only; deploy script and IAM policies in deployment/ |
| EC2 + Caddy | the live console on one instance behind Caddy for automatic TLS on an sslip.io hostname, administered through SSM with no open SSH port |
| Twilio | consent-gated outbound calls and signed inbound intake over authenticated Media Streams |
Data, and the ethics of calling people
Everything in the demonstration is a clearly labelled synthetic directory, so any
result can be reproduced. The calls are consent-gated: a call is only placed when
the exact local weekday, time, validity period, call cap, organization spacing
and permanent-suppression state all permit it. The agent discloses that it is an
automated assistant on every call, a stop request permanently suppresses future
calls, and a removal of any service is always Tier 2 or above, which means a
person decides. Audio is not persisted, and PII in transcripts is redacted before
storage. ETHICS.md in the repository states these controls in full.

Log in or sign up for Devpost to join the conversation.