Callibrate calls the authoritative source with CALL-E and corrects the record only when the transcript proves it.


Inspiration

A food pantry moves its Wednesday hours from nine-to-twelve to ten-to-one in June. The directory that thousands of people search finds out in October, when somebody takes two buses with two children and finds a locked door.

Nobody in that story did anything wrong. The pantry told the people in front of it. The directory published what it was last told. There is simply no mechanism that carries a small operational fact from the place it changed to the systems that repeat it, because the only channel that reaches a volunteer coordinator on a Tuesday afternoon is a telephone, and phoning 1.6 million service records is not a thing software could do.

United Way's 211 network handled 19 million referrals against about 1.6 million service records in 2025. Every one of those records is a claim about the world that stopped being checked the day it was entered.

CALL-E is the piece that was missing. So I built the piece that goes around it.

What it does

Callibrate is record verification by phone. When a record is doubted, it turns that record into a Verification Contract, hands the contract to CALL-E, and reads what came back.

The contract is written before the phone rings and states everything: which fields are in doubt, what counts as evidence for each, what the software may change on its own afterwards, and what ends the call. It is data, it is stored, and it is what the call is judged against.

CALL-E holds the conversation. Callibrate then does the thing that makes this a system rather than a demo: it reads the transcript itself, by rule, with no model involved. It looks for three things, in order.

  1. The provider states the value, in their own words.
  2. The agent says the whole of it back to them.
  3. The provider agrees, without hedging and without correcting.

All three, or the change does not happen automatically. A caller that reports a confirmed change the transcript does not contain gets its claim put in front of a curator with the reason it was not trusted, rather than being silently believed or silently dropped.

Then a deterministic policy grants authority. Four verdicts, and the run takes the most severe one any part of it reached.

Verdict What the call did What the software does
REFRESH The provider confirmed the published hours Moves the verification date. Changes nothing.
APPLY 09:00-12:00 became 10:00-13:00, read back and confirmed Writes it, carrying the two quotes it rests on
REVIEW "We are ending that programme next month" Leaves the record alone. A person decides.
STOP "Please take us off your list" Ends the call and suppresses the number forever

Every verification writes the contract, the CALL-E run id, the transcript, the claims and the turn ids they rest on, the verdict and its reason, and the applied diff, into an append-only SHA-256 hash chain. The database refuses updates and deletes on that table, so the only way to alter history is to break the chain, and a broken chain is visible on the front page.

The product surface a visitor sees is one button. On a listing that has not been confirmed for 91 days: Verify before I go. Press it, and a real phone call happens, and a minute later the listing either says verified by phone just now or says a person is looking at it.

The demonstration

Call one. The pantry's hours are wrong. CALL-E discloses that it is automated, asks whether nine-to-twelve is still right, hears "that changed, it is ten until one now", reads back "Wednesdays, ten in the morning until one in the afternoon", and hears "yes, ten to one on Wednesdays is correct". Three checks pass. The record changes automatically, and the two quotes travel with it.

Call two. A different provider says they are ending the programme next month. CALL-E thanks them and stops, without pressing them to confirm a closure. Callibrate blocks the automatic change, leaves the listing exactly as it was, and puts it in the decision inbox with the sentence that caused it.

How I built it

The interesting seam is that CALL-E owns the phone conversation and Callibrate owns everything about why the call happens and what it is allowed to change.

src/callibrate/calling/ is the only module that knows CALL-E exists.

  • A Model Context Protocol client for CALL-E's /mcp/openagent_oauth Streamable HTTP endpoint: brokered login with a local token cache, initialize, tools/list, then tools/call for plan_call, run_call and get_call_run. Callibrate is a server rather than an agent host, so it speaks the protocol directly instead of borrowing an agent's connection. structuredContent is preferred, and a JSON object in any text block of content is the documented fallback.
  • A contract-to-call renderer. The instruction CALL-E receives is generated from the Verification Contract, line by line: the disclosure, the current published value, the readback rule, the instruction to ask confirming questions the positive way round so a yes means yes, and the four conditions that end the call immediately. The console shows the whole thing before anybody agrees to place the call. Changing the policy changes what CALL-E is told, rather than changing a paragraph somewhere else.
  • A run that cannot be lost. run_call is asynchronous and can return before the phone rings, so the reservation row is written before CALL-E is contacted and the run_id is persisted before the first poll. A process that dies mid-call leaves a row saying a call is in flight, and the correct action is to read that run, never to place another. A run_call that returns no run_id is escalated to an operator and never retried.

Around it: a deterministic transcript reader, a four-verdict policy, a five-gate call-eligibility check, and a hash-linked ledger, in a Starlette application over SQLite with no framework in the front end.

The numbers

tests 117
labelled conversations in the evaluation 19
agreement with the label 100%
missed updates 0
false automatic mutations 0

The last row is the only one that would matter if it were wrong. A missed update costs a curator thirty seconds. A false automatic mutation publishes a wrong address for a food bank and sends somebody to it. Every threshold in the evidence rules is set for that asymmetry, and the count is on the console's front page as false_automatic_mutations, computed from curator corrections of changes the policy applied on its own.

Safety, deliberately in the way

CALL-E is integrated behind five gates, in the order they are cheapest to fail. If any refuses, no plan is created and no telephone rings.

  1. A destination allowlist, empty by default. A fresh install cannot place a live call at all. Production refuses to start with live calling and an empty allowlist. This is the gate that makes it impossible to point a demonstration at a stranger by accident.
  2. Permanent suppression after one stop request, covering every service that organization runs.
  3. Recorded consent, with a validity window, a day set, and a local-time window evaluated in the organization's own timezone. "Not before nine" means nine where the phone is.
  4. A call cap, decremented in the same transaction that reserves the call.
  5. Spacing, so a busy queue cannot phone one small charity twice in a day.

Transcripts are redacted before storage and no audio is kept. One consequence of that is stated rather than worked around: redaction masks phone numbers, so a phone readback cannot be corroborated from the record, and a phone change from a live call always reaches a person. Storing unredacted numbers to make an automatic write easier would be the wrong trade.

The default caller is a pilot line: seven scripted conversations replayed through the same normaliser, the same readback corroborator, the same reconciler and the same policy as a real CALL-E run. It dials nothing, it is labelled pilot-line in the ledger and in the console, and two of its scripts exist in order to fail. What it proves is the evidence rules, the policy, the write path and the ledger, which are the production ones either way. What it does not prove is how a real provider speaks, and no figure here is presented as if it did.

Challenges

Spoken language has to become a canonical value, or nothing downstream is comparable. "Ten until one" has to become WE 10:00-13:00, and every shortcut in that conversion is a way to publish nonsense. The reader promotes a bare closing hour by twelve only when that repairs an impossible range and leaves a plausible working day, so "ten until one" is the afternoon and "seven until seven" stays unreadable and goes to a person. Two clock times only make a range if the words between them join them, so "the one time, ten o'clock" produces nothing rather than 01:00-10:00. A day word between two times means two sessions, not a span.

A hedge in one place is a plain statement in another. The word set that makes an agreement doubtful ("actually", "but", "changed") is exactly the wrong set to apply to the turn where a provider first states a change, because "that changed, actually, it is ten until one" is not a hedge. Using one list for both sent every genuine correction to a curator until I split them.

Driving the real interface found two defects that 117 passing tests did not. The console's own content security policy was blocking every inline style the JavaScript set, and the progress stepper never left its final step, so a finished call still looked like it was ringing. Both were invisible to the API tests and obvious the moment a browser ran the thing.

Accomplishments

The architecture reduces to a diagram somebody can read in ten seconds, and every step in it is real. A trigger becomes a contract, CALL-E holds the conversation, deterministic software decides what the conversation was allowed to change, and one hash-linked row records it either way.

I am most pleased that the safety story is a property of the code rather than a paragraph in a README. The caller returns evidence and evidence authorises nothing. There is no code path where the model's account of a call reaches the database.

What I learned

The model's account of a call cannot be the sole witness to the call. That sounds obvious written down, and it is easy to violate without noticing: a single boolean flag returned by the same component that ran the conversation is enough to publish a change. Corroborating that flag against the recorded transcript is thirty lines of rules, it needs no model, and it turns a conversation into evidence.

Failing towards a person is a design parameter. The reader will sometimes miss a readback that genuinely happened. That costs a curator half a minute, and the opposite error costs somebody a wasted journey. Once the asymmetry is named, every threshold in the system has an obvious direction.

What's next

Verification Contracts are the reusable part, and the community-resource directory is one reference implementation of them. The same primitive answers "are you currently accepting new Medicaid patients", "how far out are you booking", and "is that still your opening time", and the repository builds all four contract shapes with a test for each. The last of those tests runs the same conversation under two contracts and gets APPLY from one and REVIEW from the other, because the contract changed and the domain did not.

The obvious next step is the second half of the Open Referral thesis: a provider answers once, and the corrected record publishes to every directory, agency and helpline that reads it, rather than each of them phoning separately.

Built With

Share this project:

Updates

Submission history