Inspiration

Every on-call system reports "notification sent" and treats the incident as escalated. That proves nothing. The push landed on a silenced phone. The mail went to a folder. The SMS was half-read at 03:00 and the engineer went back to sleep. The acknowledgement is the only part that matters, and it is exactly the part nobody verifies.

What it does

Ringdown walks an escalation ladder one real phone call at a time. Each rung asks one person two questions: are you taking this incident, and in how many minutes. A run ends when somebody commits with an owner and a clock, when somebody declines, or when the ladder is exhausted.

It records an acknowledgement only when the owner and the ETA are each quoted by a span the recipient actually spoke. A "yes" with no number of minutes is not an acknowledgement.

The same rules hold in Spanish. One phrase table per family carries both languages rather than one table per language, and there is no language flag, because a real on-call call code-switches and nobody knows which language it will be until somebody answers. The gates do not soften either: creo que lo tomo yo, tal vez settles as a hedge exactly the way I think I'll take it, maybe does.

"Call me back in ten minutes" is heard now, and it is still not an acknowledgement. The minutes and the words that carried them go into the ledger, and the ladder waits and rings that rung once more only if the wait plus one more call fits inside the time the ladder has left. Ask for ninety on a fifteen-minute ladder and the next rung rings now, with the request recorded rather than granted. A rung is a scope, not a person, so if the shift changed while it waited, the second call goes to whoever covers that scope now.

All of it replays in the browser with nothing to install, at https://gmassello.github.io/ringdown/ — the three scenarios step by step, and the committed ledger with a button that rewrites every verdict and reseals the entire chain. Links, seals and positions all stay green, and the verification fails anyway.

That one case is the whole product. In the demo's second scenario the provider is perfectly satisfied — the call completed, task_completed is true, confidence is high at 0.91, the disposition reads acknowledged — and there is no ETA. A system that branches on those three signals reports the incident as escalated and goes back to sleep. Ringdown drops a rung, and the backup commits.

An alert arrives through a mapping file, not through vendor code. PagerDuty and Opsgenie both land as incidents without a line in the adapter that knows either name — Opsgenie is the one that tests the claim, because its payload carries no priority, no description and no link back, and the mapping file absorbs all three. When the verdict settles, the run writes it back to PagerDuty as a note quoting what the engineer actually said. A note, never an acknowledgement: suppressing a pager’s own escalation on the strength of a phone call is the operator’s decision.

A model may write that mapping file, and that is the only place a model touches anything. suggest-mapping asks Gemini to draft the mapping for a payload nobody has mapped yet, then runs the draft: the adapter executes it and the incident loader validates the result before it reaches disk, sending a rejection back once with the loader’s own words attached. The deterministic side corrects the model, never the other way round. Nothing in preview, run or verify reaches for one, and a layering test asserts that only the CLI can even import it.

The same engine also chases something that is not an incident at all: what the agent says is a template the incident file can replace, so the ladder, the verification and the ledger pursue a supplier over a missed service level with no code that knows what an SLA is.

How we built it

It places the call over REST and re-reads it over MCP. The provider's two surfaces are not two views of one JSON document: REST reports lowercase statuses and exposes task_completed and completion_confidence; MCP reports uppercase statuses and accepts no extraction schema at all. So verifying over MCP re-derives the acknowledgement from a transcript served by the transport that did not write it. An agent that audits itself through the channel it wrote with proves nothing.

Ten checks run on the attempt that acknowledged, in two blocks that prove different things: six establish that both surfaces describe one call, four re-derive the acknowledgement from the second channel's transcript. The split matters because the second group re-runs Ringdown's own extractor: it catches a transcript that differs between the two surfaces, not an extractor that read one transcript wrong. Buying that would take a second derivation the provider cannot supply. Zero checks is not success — a ladder with no attempts is never reported as verified.

Every verdict is sealed in a hash-chained JSONL ledger that does something a flat append-only log cannot: it re-derives the verdict from the recorded attempts on replay. Rewrite a verdict, recompute its hash, relink every record after it, and the chain closes cleanly — and verification still fails.

Python 3.11 with an empty dependency list. 606 tests that run with no credentials and no outbound calls, against a fake CALL-E that serves both surfaces from one store with two different projections.

There is a second app, calle-receiver, that exists for one reason: CALL-E does not dial Argentina. It is a FastAPI service that answers the agent on a US Twilio number and bridges to an Argentine phone, with recording, live transcription and a dashboard.

Challenges we ran into

The idempotency key stopped being a nicety and became the load-bearing part. Every REST create times out before the provider answers: five for five in August, three for three in September, eight of eight across two sessions three weeks apart, every one of them at a 15-second client socket timeout — one create sat inside the provider for four minutes and twenty-seven seconds before it did anything. Every one of them had in fact created the call. Because the key is derived from the payload rather than generated per run, every replay returned the existing call instead of dialling a second time: eight creates, eight replays, and nobody was dialled twice. Against the real API, reconciliation is the normal path, not a rare branch.

Speech recognition turned out to be the honest adversary. On one connected call the provider reported task_completed: true at 0.86 high, with evidence reading "The engineer acknowledged taking the incident." The recipient's turn behind that reads "Yes. I'm banking this incident right now" — the transcriber heard "banking" for "taking". Ringdown refused the acknowledgement, correctly: no commitment phrase was ever spoken in the transcript it was given.

Three weeks later we placed three more calls and found three things worse than a misheard word. The agent does not wait for an answer it was told to wait for: our task says, in these words, do not describe the incident until they have answered "Am I speaking with {name}?" — and it read the incident straight through the question, before the recipient had said anything. Non-English speech comes back as English phonetics: Sí, soy German was transcribed as "C is not a" and Sí, lo tomo yo as "C, the Thomas" — not low-confidence guesses but confident English words, on a call the provider's own completion_confidence called high at 0.9, and CreateCallRequest exposes no language parameter at all. And on that same call the agent told the recipient it had recorded the acknowledgement while the API returned task_completed: false. Ringdown fails closed on that flag, which is right, but the person hung up believing they were on the hook.

Four of the first six calls ended three seconds in, reported as the recipient hanging up, while the Twilio account that owns the destination number had no record of any of them — the phone never rang. Three of three connected in September, so we stopped claiming it still happens and said so instead. The attempt now carries its measured duration, the reason is recorded whatever status the provider reports, and the run prints how many calls ended that way beside the verdict: a ladder can no longer report an incident as unowned when no telephone ever rang.

Accomplishments that we're proud of

The demo output was written by hand before the code that produces it, and a test now compares every quoted block against a real run, so the contract cannot drift silently.

And we placed nine real calls against the live provider, six on 20 August and three more on 13 September — which the repository itself called "the single highest-value thing anyone with a dialable number can do".

The demo video is not a reconstruction either. Its last fifty seconds are that August call: the phone ringing, the call answered, and the words the provider transcribed appearing beside it at their own offsets — including the moment it heard the wrong name and had to ask again.

And the contribution was reviewed and merged into CALLE-AI/awesome-phone-call-agents — the skill and the bridge now live in CALL-E's own repository. A second pull request, #512, brings that copy up to date with everything above: 55 files, and the shared repository's own validator passes.

What we learned

Those calls confirmed the REST contract: metadata echoes back exactly as sent, so an attempt can be tied to the call it belongs to.

They also settled an open question, in the direction we did not want. get_call_run takes a run_id that only run_call hands out. No identifier a REST-placed call exposes resolves to one — not the call id, not provider_call_id, not the attempt or recipient id. And the first successful run ever observed does not have the shape our parser reads. Cross-surface verification cannot be completed from this side today, so a live run reports the acknowledgement as unconfirmed rather than claiming it — exit 45, which means the second channel said nothing, never the second channel disagreed.

We did not hide that. It is stated as a known ceiling, the provider's actual responses are committed as test fixtures with their provenance, and one test asserts that the parser cannot read a real run — so the gap is a tested fact rather than a surprise waiting for the next person. The same findings went to the CALL-E team as feedback.

What's next for Ringdown

The provider fix is small and unlocks the whole category: have get_call_run accept a call id as well as a run id, or expose the run id on the CallTask that REST already returns. Either one makes an agent auditable across transports.

On our side: re-escalation when an ETA expires, and a keyed HMAC over the ledger so it proves something against an adversary and not only against accident.

Built With

Share this project:

Updates

Submission history