The problem

Before a utility switches off power to stop a wildfire, it owes notice to customers on the Medical Baseline programme — the households that depend on powered medical equipment. That duty is not satisfied by sending a notice. The notice has to be received.

A voicemail is not receipt. A text is not receipt. A call that rang out is not receipt. PG&E states that Medical Baseline customers must confirm receipt, and that if they do not respond it will keep trying hourly or contact them in person until they are reached.

The failure here is quiet. A call reaches an answering machine, the row gets marked notified, and nobody visits. The person finds out when the lights go off.

What it does

PositiveContact takes an event and a customer roster and runs the whole outreach loop.

It validates the roster and refuses rather than repairs: a phone number not already in E.164 form is a blocking error, not something to fix by guessing a country code. A timezone comes from the roster, never from the phone number. A language the destination line does not support routes to a bilingual human callback instead of being called in English anyway.

Every call is disclosed before anything else — who is calling, that it is automated, that it may be recorded — and it asks for nothing: no account number, no payment, no date of birth. Utility impersonation scams are common, so the real call is built to be easy to tell apart from the fake one.

A contact is confirmed automatically only when the transcript shows a human acknowledged the notice. Four things have to hold. CALL-E's structured result has to say a live person answered and acknowledged. A second reader has to find that acknowledgement in the transcript independently, in a turn after the notice was read out. And completion confidence has to pass the gate — high, or a score at or above 0.80, where missing or unrecognised fails. If any of the four fail, it is not a confirmation. The only other route is a named operator confirming with cited evidence, and the report counts those separately.

Everything else moves the escalation ladder forward: retry the primary number, try the alternate contact, retry again, and when the ladder or the clock runs out, prepare a field visit for a named person to approve. Wrong numbers and refusals are never dialled again. A contact in human review is still on the clock — unresolved at the cutoff, it becomes a field visit automatically.

If a resident asks for help with a prescription or powered equipment, the agent gives no medical advice. It records a coded route, broad timing, whether permission was given to contact a provider, and whether immediate danger was reported. An operator can then authorise one call to an approved pharmacy or supplier that asks only about general availability and carries no resident name, number, address or clinical detail.

Every count in the report carries the denominator it was measured against:

Metric Value
Enrolled contacts in scope 12
Contacts attempted 11 of 12 in scope
Live human reached 9 of 11 attempted
Positive contact confirmed 7 of 9 live reached
Voicemail only 1 of 11 attempted
Wrong number 1 of 11 attempted
Refused 1 of 9 live reached
Language unsupported, callback opened 1 of 12, never dialled
Field visits pending approval 1
Calls placed 16

Sixteen calls were placed and seven people were confirmed. A report that printed only the first number would be true and useless.

How it was built

Python, FastAPI, SQLite, and the CALL-E Calls API over REST.

The application owns business state, CALL-E owns call state, and the two never get collapsed. Every call starts as a durable intent written down before anything is sent, with an idempotency key derived from what was authorised — event, contact, ladder step, target — and never from the attempt. A retry replays the same key; a restart re-derives it rather than minting a new one.

Webhooks are a wake-up signal and nothing more, because the published contract has no authentication on them. The receiver validates the envelope, writes one row, and returns. A worker then re-reads the call from the authenticated API and checks it is bound to the intent we authorised before anything acts on it.

State changes come from one table of allowed transitions; anything else raises. The audit trail is append-only, enforced by database triggers rather than convention. The raw phone number lives in exactly one place and everything else gets a masked form — a test walks every table and fails the build if a raw number appears anywhere else. 459 tests, all offline, with a fixture that fails any test opening a network connection.

Two things that nearly shipped wrong

The platform cannot tell you an answering machine picked up. Nothing in the published contract carries that signal, and confidence is not a proxy for it — another submission published a voicemail at 0.93 high sitting next to a real success at 0.94 high. So detection had to move into reading the transcript for greeting patterns, which is a per-language lexicon, and a greeting outside it reads as a person. Ours let one through: "Hi, yes, this is the Smith family, we're not home right now, please leave your name and number." Because it contains "yes", the call adjudicated as a live person acknowledging, with the machine's own greeting stored as the evidence. Found in review, lexicon widened, test written.

A denial written with an apostrophe was invisible. The negation pattern was \bn'?t\b, which looks like it catches "didn't" and cannot — the character before the "n" is a word character, so \b has no boundary to match. "I couldn't hear a word. Correct?" adjudicated as confirmed, quoting the denial as its own supporting evidence. A contraction and its expansion are now asserted to reach the same outcome.

What it does not do

There is no cancel. Once CALL-E returns a call id this application cannot stop that call, because the API publishes no cancel endpoint — there is deliberately no cancel method on the transport interface and a test asserts its absence.

Field visits are prepared, never dispatched; a named person approves each one. No medical advice, ever: no field records a condition, diagnosis, medicine or equipment model, and a redactor strips health words from free text on the way in.

Filed upstream

Three findings went to the CALL-E team with a proposed API shape, rather than staying notes in our own repo:

  • #477 — the result-schema constraints are split across the two schema fields, and maxLength is named in neither the supported nor the unsupported list, leaving three failure modes an implementer cannot tell apart.

  • #478 — proposes a start_at scheduling field on POST /v1/calls, so every integration with a time dimension stops having to write its own scheduler. For a workflow built on quiet hours, the failure mode is a phone ringing in somebody's night.

  • Comment on #406 — proposes an answered_by enum, so an application can tell afterwards whether a person was ever on the line.

Try it yourself

No credentials needed, and it cannot place a call.

git clone https://github.com/Abhinav0905/CALLE-AI.git
cd CALLE-AI && git checkout feat/positive-contact
cd apps/python/positive-contact
python3 -m venv .venv && ./.venv/bin/pip install -e ".[dev]"

./.venv/bin/pc preflight                                  # masked plan, no calls
./.venv/bin/pc run --mode fixture --stop-before-cutoff    # the ladder, held at the cutoff
./.venv/bin/pc serve                                      # dashboard on 127.0.0.1:8000
./.venv/bin/pytest -q                                     # 459 tests, offline

A real call needs three things a person has to supply: PC_MODE=live, the --i-understand-this-places-real-calls flag, and an explicit --max-calls N. It then prints the masked plan and the exact words the call will use, and waits for you to type PLACE.

The live demo is read-only — operator actions are disabled unless it is run locally.

Built With

Share this project:

Updates

Submission history