On the Line

On the Line is a community resource directory that phones ahead. It calls the food banks, clinics and shelters it lists, confirms what is still true, publishes only what a provider said out loud and confirmed on readback, and hands everything that needs judgment to a person. When a neighbour reports a closed door, that report becomes the next call in the same database transaction.

Check it while you read

The deployment is live: https://54-84-108-211.sslip.io

  • The public side needs no account. https://54-84-108-211.sslip.io/help is the search a person looking for help would use. Take a referral, then report a closed door and watch that report reach the top of the call queue.
  • Sign in to the curator console with judge / ontheline-demo-2026.
  • On the overview, press Call the provider on the first queued service and listen to the call. Turn the sound on.
  • Decisions holds the closure the agent refused to publish, with the evidence diff and the provider's own quote.
  • Activity is the append-only chain the database will not let anyone edit or delete.
  • Add ?demo=1 to any page for a narrated walkthrough that drives the real console. Escape stops it and hands the product back.

The problem, in one trip

Two buses. Two children. A pantry that closed in July and a directory that still says Wednesdays, nine to twelve. The referral was free to send and expensive to follow.

United Way reports that the US 211 network made 19 million referrals across 1.6 million locally curated services in 2025. Open Referral documents the maintenance loop behind that number: fragmented directories repeatedly ask the same service providers for the same updates, staff spend their time re-verifying records, and the people relying on the listing get less reliable information. Keeping those listings true is repetitive work that almost nobody is funded to do, so it does not get done, and the cost lands on the person least able to absorb it.

Who it is for, and why it matters

The same phone call helps three different people, which is what makes this a Good Neighbor project rather than a data-quality tool.

Who What they get
The person looking for help A listing a human being confirmed, with the date they said it
The volunteer holding the directory together The routine calls made without them; only judgment reaches their inbox
The organization being listed Called once instead of twelve times, because the answer is published as open HSDS any other directory can consume

That third row is the part most directory tools leave on the table. Open Referral's own strategic case is that fragmented directories pester the same service staff for the same updates. On the Line exports stable-UUID HSDS 3.2 on a public endpoint, so the result of one call is publishable to every other directory. Call once, publish for everyone.

The mechanic

Someone finds the door locked and says so. In the same database transaction, that service jumps to priority 1000 at the top of the call queue, and the agent phones the provider. The next neighbour finds it open.

A neighbour reports a closed door, and it becomes the next call

Help-seekers are the freshness sensors. The provider calls are the repair crew. A directory chatbot does not close this loop, because it answers the question it was asked and then forgets it happened. Here the failed referral and the new verification task share one transaction, and a direct failure is merged into any existing open task for that service, so the queue never double-counts.

What the call sounds like

The console places the call and you hear it. The agent opens by saying what it is, reads the changed hours back before recording them, and waits to be told it got them right:

On the Line: Let me read that back. Wednesdays, ten in the morning to one in the afternoon. Have I got that right?

Provider: Yes, ten to one on Wednesdays is correct.

That readback is the evidentiary standard, not decoration. Without an explicit statement and a confirmed readback, the change is held for a person. The audio on screen is rendered turn by turn from CALL_SCRIPTS, each speaker in a separate voice through a 300 to 3400 Hz telephone band, and the transcript reveals each line exactly when it is spoken, so the words you read and the words you hear cannot drift apart.

A verification call, streamed turn by turn: the disclosure, the readback, and the change stated

How Strands is used

The agents do the interpreting; plain Python decides what gets written.

Voice, both directions: a Strands BidiAgent on BidiNovaSonicModel runs the provider verification calls and the inbound help-seeker line, with interruption handling. src/ontheline/audio/telephony_io.py bridges a carrier stream to the agent: mu-law 8 kHz in from the phone, PCM to the agent, PCM back, mu-law out, with bounded backpressure so a fast talker cannot outrun the buffer. On a configured deployment the carrier is Twilio bidirectional Media Streams; every HTTP callback and WebSocket upgrade is verified against Twilio's signature over the canonical public URL, and audio is never persisted.

Extraction: a GraphBuilder graph fixes the order, a safety reader then a structured extractor, entered only after a call has produced a transcript. The safety reader preserves every provider statement verbatim and flags ambiguity, distress, a help-seeker answering instead of staff, and stop requests. The extractor runs on Amazon Nova 2 Lite with structured_output_model=ExtractionResult, so its output validates against a Pydantic schema before anything downstream sees it.

A domain contract, not just a JSON shape: the extractor carries a Strands GoalLoop. JSON validity is table stakes; the goal validator enforces domain invariants and retries the model until they hold. A non-active status has to set removes_service=true. A read_back_confirmed change has to also be explicitly_confirmed. A model that returns well-formed but domain-invalid output is sent back with specific feedback rather than accepted.

Lifecycle and state: invocation hooks (HookRegistry, BeforeInvocationEvent, AfterInvocationEvent) record which agent ran and in what order, and the trace is persisted with the verification. FileSessionManager keys graph state on a path-safe, task-specific session id, so a retry resumes rather than starting over.

AgentCore: Dockerfile.agentcore builds a stateless ARM64 runtime exposing /ping and /invocations. It runs the extraction graph and returns only schema-valid extraction plus the lifecycle trace. It holds no directory database and its IAM role grants Nova invocation and telemetry only, so policy, consent and persistence stay in the application boundary. OTL_AGENT_MODE=agentcore selects it; bedrock runs the same graph in-process; demo runs a deterministic offline adapter that escalates anything it cannot establish explicitly.

Evaluation: the official Strands Evals SDK runs a versioned experiment of 12 adversarial provider personas. The committed deterministic report records 100% exact policy matches, a 0% false-apply rate, a 0% over-escalation rate, and successful blocking of every tested removal, missing disclosure and stop request.

The line the model never crosses

Model output is a proposal. A deterministic Python policy in ontheline.policy classifies complete evidence into one of four tiers before any transaction touches the directory.

Tier Evidence What happens
0 Every requested field confirmed unchanged Refresh the assurance date
1 Allowlisted field, explicit statement, confirmed readback, canonical value Apply it, with the quote attached
2 Removal, eligibility, ambiguity, contradiction, low confidence Hold it for a person
3 Missing AI disclosure, distress, help-seeker reached, hostility, stop request End the call, flag, and suppress when requested

The policy validates field coverage, confidence, explicit confirmation and readback, canonical HSDS schedules and E.164 phone numbers, high-judgment fields, and removals. The model is never handed a database, arbitrary SQL, a publish tool, a delete tool, or consent authority. Changes are atomic per verification: the overall tier is the maximum across every proposed change, so if one field needs judgment, nothing else from that call is applied on its own.

The evidence chain

Every automatic change and every human decision keeps its call, timestamp, quote, consent reference, policy tier and agent trace in a SHA-256 hash chain. Each hash covers the prior hash, the canonical payload, the event type, the actor and the timestamp, and database triggers reject every UPDATE and DELETE on the audit table. The human-interrupt requirement is a database-backed boundary rather than a pause inside a model process, which is what lets a pending decision survive a restart and be resolved by an approve, edit or reject transaction that revalidates the field before it publishes.

What the build settled

A readback is the input, not a formality. The first version of the call had a two-turn transcript with no readback in it, while the policy required one and the tier table advertised one. Making the call a real conversation, where the agent reads the changed value back and waits to be told it is right, is what lets a hedged or missing confirmation land in the decision inbox instead of the directory. The conversation is the evidence.

A goal is a stronger contract than a schema. Nova 2 Lite returns valid JSON readily. The failures that matter are domain failures: a closure that does not mark the service for removal, a readback flag set without an explicit statement behind it. Encoding those as a Strands GoalLoop goal, with feedback the model retries against, is what turns a schema-valid answer into a domain-valid one.

The interrupt belongs in the database. Putting the human gate in SQLite, not in the interface, is what makes "only judgment reaches a person" true across a restart, a redeploy and a second worker. An interrupt that lives in a process dies with the process.

A guard that has never been fired once is an assumption. The live site restores its demonstration directory on a timer. Testing that timer once, on purpose, before letting it run unattended during judging, is the only thing that could have caught what it was doing: an early version tore down the stack with its volumes and discarded the Let's Encrypt certificate on every reseed, which would have exhausted the certificate quota and taken the site down mid-judging while every containers-came-back check still passed. It now restores only the data volume and leaves the certificate in place.

Technologies used

AWS and Strands

Strands Agents SDK BidiAgent on BidiNovaSonicModel for two-way voice, a GraphBuilder extraction graph, structured_output_model, a GoalLoop domain validator, HookRegistry lifecycle tracing, FileSessionManager resumable sessions, and the Strands Evals SDK
Amazon Nova 2 Sonic live bidirectional speech for the provider and help-seeker conversations
Amazon Nova 2 Lite the structured HSDS extractor behind the graph
Amazon Bedrock AgentCore a stateless ARM64 extraction runtime (/ping, /invocations) that returns schema-valid extraction and trace only; deploy script and IAM policies in deployment/
EC2 + Caddy the live console on one instance behind Caddy for automatic TLS on an sslip.io hostname, administered through SSM with no open SSH port
Twilio consent-gated outbound calls and signed inbound intake over authenticated Media Streams

Data, and the ethics of calling people

Everything in the demonstration is a clearly labelled synthetic directory, so any result can be reproduced. The calls are consent-gated: a call is only placed when the exact local weekday, time, validity period, call cap, organization spacing and permanent-suppression state all permit it. The agent discloses that it is an automated assistant on every call, a stop request permanently suppresses future calls, and a removal of any service is always Tier 2 or above, which means a person decides. Audio is not persisted, and PII in transcripts is redacted before storage. ETHICS.md in the repository states these controls in full.

Share this project:

Updates

Submission history