Inspiration

Every business that sends invoices has someone whose week is spent on the phone asking when a bill is going to be paid. It is the least automated corner of finance, and not because the calls are hard — they are short and repetitive. It is because the output of the call is soft. Someone says "it's in the next payment run," a colleague types "will pay soon" into a CRM note, and four weeks later nobody can tell you whether that was true, which customers say it every single month, or what the company would gain by chasing differently.

We started from a blunt observation: a phone call cannot collect money. Everything it produces is a sentence, and roughly half of those sentences are not kept. So the interesting object is not the call. It is the promise.

What it does

PromiseLedger reads an AR aging file, decides which overdue invoices are eligible to call today, places one CALL-E call per invoice, and turns each result into one of a small set of honest outcomes.

When someone commits to a date, that becomes a promise: a durable record carrying the amount, the date, the intent that authorised the call, and the exact words from the transcript that support it. Promises expire. When the date passes and no payment arrived, the promise is broken — and that, not an attempt counter, is what escalates the next call from a friendly reminder, to a follow-up on the missed date, to a request for the AP manager.

Three design decisions carry the whole project:

1. A promise needs a receipt. CALL-E returns an excellent structured result, but that result is a model's reading of a conversation. We never let it drive a business transition alone. Before a promise is recorded, the date must appear in a transcript turn spoken by the counterparty — as an ISO date, "the 5th of September", "the 5th", "next Friday", "end of the month" — and the amount must be named or clearly stated as the full invoice. "It should go out soon" fails corroboration and becomes needs_human. That single rule is the difference between a ledger of commitments and a ledger of wishful thinking.

2. Escalation on promise history, not attempt count. An attempt counter treats a customer nobody has reached exactly like one who has broken two commitments. Those two people need completely different phone calls. Ours get them.

3. The call is structurally prevented from becoming a payment channel. On a call about money, people offer to pay. An AI agent must not accept: it cannot verify who answered, and a card number spoken aloud is instantly copied into a transcript, a summary, an audit row, and a model context window. So the task text tells the caller to interrupt and redirect to the remittance portal — and because that layer will sometimes fail, a firewall scans every string arriving from CALL-E before anything touches disk, redacting Luhn-valid card numbers, IBANs, UPI IDs, routing and account numbers, and credential shapes. Detection does not merely redact: it fails the call closed, discarding even a perfectly good promise, because a call where someone tried to pay is a call a human should read.

Around all of that sit the parts that make it safe to point at real customers: calling windows enforced in the recipient's own timezone, hard call caps, permanent suppression on request, a dispute that stops the ladder rather than accelerating it, and an HTML console reporting outstanding balance, promised-but-not-yet-due, and the kept-versus-broken rate.

How we built it

Python 3.12, the official calle-ai SDK, SQLite for durable state, and nothing else.

We followed the repository's own docs/production-workflows.md as a specification. The business intent is written durably before the call boundary. The idempotency key is derived from the authorization — invoice, cycle, authorization id, task version, schema version — rather than from the attempt, so a crash or a reconnect cannot become a second phone call. An ambiguous submission lands in submission_unknown and blocks the next cycle instead of redialing. Every terminal snapshot is verified against the reserved intent — call id, metadata, task and schema version, and the exact approved destination — before anything is believed. The webhook receiver commits one inbox row and acknowledges, treating an unsigned delivery as a wake-up signal that is always reconciled against an authenticated read.

Twelve scripted scenarios drive the complete path — reserve, submit, poll, verify, scrub, classify, apply — with no network and no credits. That is how we developed every branch of the result policy before spending a single call.

Challenges we ran into

The spoken-date problem was harder than expected and is where most of the interesting work went. People do not say "2026-09-05". They say "the 5th", "next Friday", "end of the month", "in a fortnight". We resolve a deliberately small, well-tested set of those forms against the call date, and everything outside it fails to needs_human. Being narrow beat being clever: the system is allowed to say it does not know.

The second was resisting the pull to let the call close the loop. It is very tempting to mark an invoice paid when the customer says they paid it. We do not. A phone call has no visibility of a bank account. That result is paid_claim_unverified, and finance matches the receipt.

The third we found late, while building the interactive trace page. A call that simply rang out was being classified needs_human — it had no result, so result validation failed before the terminal-status check ever ran. That is wrong in a way that matters: routing every no-answer to a person buries the outcomes that genuinely need reading, which is the exact failure mode this app exists to prevent. Reordering the two checks fixed it, and two regression tests now pin the behaviour.

Accomplishments that we're proud of

77 tests that run offline with no credentials and place no calls. They cover the firewall's true and false positives, every spoken-date form the corroborator accepts and rejects, each rung of the escalation ladder, the fail-closed dispositions, the crash and ambiguity paths, ten concurrent webhook deliveries of one event against the real HTTP server, and a contract check that fails if an SDK upgrade renames a field we send.

The firewall proof is the one we like most. In the recorded card-number scenario, the digits never enter the stored snapshot at all — so they cannot reach a log, an audit row, or the public demo page built from that data. And no promise was written, even though that same call also produced a valid one.

What we learned

Reading the CALL-E API design taught us most of this. The result-schema documentation explicitly recommends string enums over booleans and an unknown value for anything the call may not establish — an opinionated choice, and the right one. completion_confidence arriving alongside the result is what let us put a confidence floor under promise recording. Once we had those two, the shape of the whole app followed: let the provider report what happened, and keep every judgement about what it means on our side of the line.

We also learned how much of a phone workflow is deciding not to call. Grace periods, calling windows, caps, cooling-off gaps, open promises, disputes, suppression — on a typical day most invoices in the sample are ineligible, and that is the system working.

What's next for PromiseLedger

Kept-rate as a per-customer credit signal. A promise-weighted cash forecast, where each open promise is discounted by that customer's historical keep rate. And ERP connectors, so the aging file and the payment feed stop being CSVs.

Built With

Share this project:

Updates

Submission history