Inspiration

A medical biller spends about 25 minutes on one claim-status call: the phone tree, the hold, one question, one answer copied into a financial record. The 2024 CAQH Index puts the industry's bill for that at $11 billion a year. CALL-E can make the call. What nobody had answered was what happens after the call, when the agent comes back and says "Claim 4471 was paid $1,240 on August 12." That sentence is perfectly shaped whether it is true, invented, or collected from the wrong department. A biller cannot tell the difference, and the amount goes into the ledger either way.

I wanted to build the thing that makes the caller prove it.

What it does

Kol makes CALL-E prove a payer call. Before a claim-status answer is accepted, four witnesses have to agree, and none of them can be written by the model:

  1. Question witness. The claim-specific question appears on the agent side of the transcript.
  2. Answer witness. The payer's exact words contain the claim reference and the status, amount, date or denial code.
  3. Destination witness. A payer quote establishes that the claims desk answered, not provider services with figures that happen to agree.
  4. Route witness. An account of the keypad route that is independent of the model: a fixture log, decoded DTMF audio, a provider event, or the person who answered.

Missing evidence is held for review. Conflicting evidence is contradicted and never written. Only a verified call may teach the Route Atlas, which remembers the phone tree so the next chase replays it by keypad instead of exploring. If the tree moves, the route is quarantined rather than trusted.

Beyond one call, one claim:

  • Several claims on one call. A biller reads six claim numbers off a list once they are through the tree. Kol makes the representative name each claim with its answer and binds every answer to its claim. That catches the failure only a batch can have: every number spoken, every quote real, claim A filed with claim B's amount. The single-claim gate passes it. The batch gate contradicts it.
  • CALL-E's own evidence is cross-examined. The provider's free-text justifications are the model explaining itself. One that narrates an amount the payer never said blocks auto-accept.
  • A PHI guard that refuses to dial. A member ID, date of birth, diagnosis code, patient name, SSN, email or phone number in the request stops the call before it is compiled, and the refusal masks what it found.
  • Every live chase is a durable run. The call is created exactly once, the run suspends until the person who answered says which tones they heard, and it never dials twice.

How I built it

Only possible because one CALL-E request returns three things together: a speaker-labelled, time-offset transcript, a strict structured result, and keypad navigation of the phone tree driven from task prose. Kol turns those into evidence that can disagree with each other.

  • The Route Atlas compiles a stored route into the task text, since the API has no navigation parameter, and reads the route back from the structured result.
  • I wrote a deterministic TypeScript verifier with zero runtime dependencies that binds every field to a payer-side transcript span. It can only reduce trust, never raise it.
  • Each live chase is a Vercel Workflow run. The create step runs once, with retries disabled and an idempotency key derived from the run; polling is checkpointed; after the call ends the run suspends on a hook until the operator's receipt resumes it, a minute or a day later. The verdict is the run's return value. A crash test kills the server mid-run, restarts it, and proves the same run finishes verified with exactly one CALL-E create.
  • The judge-facing product is Next.js on Vercel: a landing page, an evidence console, Route Atlas drift views, and a guarded live lab with three CALL-E modes, including an IVR route replay where the person who answers reads a two-level menu aloud, CALL-E replays the route by keypad, and the operator attaches the tones they heard as a receipt the model cannot write.
  • The judge-facing product is Next.js on Vercel: a landing page, an evidence console, Route Atlas drift views, and a guarded live lab with three CALL-E modes. In the claim-evidence mode, a real call's transcript and structured result go through the witness gate on screen. The IVR route replay mode is proven end to end against a stand-in CALL-E: in three live calls where I read the menu aloud, CALL-E's agent talked over it and pressed no keys, so live keypad replay still needs a real automated IVR.

Challenges I ran into

The public API has no independent keypress witness. The model reports its own keys, and a model grading itself is not evidence. I refused to paper over that: a live result stays under review until a fixture log, decoded audio, a provider event, or the person who answered corroborates the route. That constraint became the product's safety boundary, and the live lab shows it on screen: the model's account is labelled testimony, and the receipt comes from someone else.

Building a free IVR fixture from India was its own fight. Amazon Connect and Chime refuse Indian-entity accounts, and Twilio trial numbers reject any caller they have not verified. So I built a Goertzel DTMF decoder and let the laptop play the menu into a phone on speaker.

Shooting the demo found two real bugs: a polling race that hid the live transcript on fast connections, and an unreachable "run started" panel. I fixed and deployed both.

The live test taught me the most. In three calls where I read a phone menu aloud, CALL-E's agent talked over it and never pressed a key, even when told to stay silent and use only the keypad, so live keypad replay is still unproven and I say so. The conversational call worked: fourteen turns, the agent read the year back wrong and asked me to clarify, and CALL-E returned a correct structured result with 0.90 confidence. Kol still marked it contradicted, because the phone transcript heard the claim number as "104471" and the extraction never named the department. That is the product in one call: a correct answer without its evidence does not get written.

Accomplishments that I'm proud of

  • 108 automated tests, run in CI on every push.
  • 2,000 adversarial verdicts, regenerated from code, with zero unsafe auto-accepts: 720 single-claim cases across nine failure families, and 1,280 claim verdicts on multi-claim calls across five.
  • A test that proves the single-claim gate passes a swapped answer and the batch gate catches it.
  • A durable chase that survives a server crash mid-call with exactly one CALL-E create.
  • A merged contribution to CALL-E's awesome-phone-call-agents: the kol-ivr-route skill and the apps/typescript/kol app.
  • A public demo any judge can open without spending a call.
  • A real CALL-E call, shown end to end in the demo, where Kol held a correct-looking answer as contradicted because its evidence was not in the transcript.

What I learned

Structured output is an interface, not evidence. A confidence score can lower trust but must never raise it. A verifier that redacts everything is its own kind of failure, so every check has a clean case it must accept. And a long-running call with a human at the end of it is a workflow, not a request: it has to outlive the tab, the function, and the day.

What's next for Kol

Interview revenue-cycle operators to measure how often confident wrong answers reach the ledger today. Connect a controlled IVR fixture for fixture-logged route receipts on every live call. Measure per-payer false-accept and review rates. Design a compliant pilot with counsel: BAAs, access controls, retention, human review.

Built With

Share this project:

Updates

Submission history