Inspiration
My wife is a cardiac rehab nurse. The problem she describes is not that patients refuse treatment — it is that they quietly stop coming, and nobody notices until the file is closed months later.
Cardiac rehab is a course: eight, twelve, twenty sessions over weeks. Missing one session means almost nothing. Missing your second-ever session means something quite different from missing your seventh after six good weeks. The clinic sees the same row in the same diary either way.
So the interesting problem was never "call the patient". It was deciding which patient to call, and which conversation to have with them.
What it does
Every other phone-agent I found makes one call about one event. This one decides across a course.
A clinician defines the course, the bookable slots, and — importantly — a ladder of things the clinic authorises the caller to offer. Then, for each patient:
- Read their attendance history and classify where they sit on their trajectory through the course. Twelve states, checked in a fixed order.
- Run a gate that can refuse to dial and says exactly why —
consent_missing,phone_not_e164,cooldown_active,attempts_exhausted,promise_window_open,no_offered_slots. - Build a call goal from a whitelist of validated fields, plus a closed result schema whose slot enum is generated from that course's live data.
- Place one call through CALL-E, serially, with an idempotency key derived from the request itself.
- Read the result back conservatively, record what the call cost, and put one recommended next step in front of the clinician.
The differential, which is the whole point:
1. p_ivy -> first_slip / call_light_rebook
why: one missed session after 6 attended; rebook without interrogating
2. p_omar -> early_slip / call_blocker_and_rebook
why: missed a session with only 1 attended; early drop-off risk is highest here
Same missed session, same date. Ivy's call books her back in and is explicitly told not to ask why she missed it — she has attended six times and does not need interrogating. Omar has been once, which is where a course is actually lost, so his call asks what is stopping him. A regression test asserts these differ; if it ever passes trivially, the app has become a reminder bot.
It negotiates, but only with permission. If neither offered time works, the caller may spend only what the clinic authorised, in tier order, one at a time: an evening group, then a taxi voucher, then a phone review, then pausing the course. It cannot invent an offer. And the ledger records which tier it had to spend, because a booking that cost a taxi voucher is not the same result as one that cost nothing.
It knows what it must not do. A volunteered symptom is never assessed, never reassured, and never rebooked — even when the same call also accepted a slot, because the symptom check runs before the booking check. A voicemail is told only who the message is for and which number to ring; nothing about the programme, the specialty, or a missed session.
How we built it
- Python 3.11, standard library only for the core. ~1,780 lines.
calle-ai==0.7.0behind an optionalliveextra, so a default install cannot place a call —calleis not merely unused, it is absent.- CALL-E one-shot Calls API:
calls.create(task, recipients, result_schema, metadata, idempotency_key)→calls.wait_for_result(...).
The architecture has exactly one seam. workflow.py depends on a
typing.Protocol called CallPort, never on CALL-E. Two implementations satisfy
it — FixturePort (a JSON file, no network) and LivePort (the SDK) — so the
test suite, the demo and the live run are the same code path, with no
if live: branch anywhere. Every untested branch is one that runs for the first
time on demo day.
Because the model behind a CALL-E call is not selectable, the only two levers are
the goal string and the result schema. Both are used narrowly and deliberately:
the schema is closed (additionalProperties: false), almost every field is an
enum, and chosen_slot_id is enumerated from that course's live slots, so a
slot the clinician never offered is unrepresentable in the output. That is
stronger than validating afterwards — the wrong answer cannot be expressed.
Why one-shot Calls rather than Goals. CALL-E's docs draw the line: "Use the one-shot Calls API when each request needs new task text or a request-scoped result schema." This app composes a different script per patient from their history, and builds the slot enum from live course data. Both conditions apply.
Challenges we ran into
The first idea died in week one, and that was the best outcome available.
The original plan was a general "phone canvass" skill. An audit of the live repo
found CALL-E's own outbound-call-skill-creator already generates it from a CSV
— and that two other entrants had filed the same idea four days earlier. The
planning method had been "read merged main, compare against the roadmap", which
is structurally blind to an open PR queue. In a hackathon whose deliverable is a
PR to a shared repo, the open queue is the competitive field. The
disqualifying evidence cost one gh pr list. It had been scheduled third.
Then five real calls found six faults that no fixture could have.
| # | What broke |
|---|---|
| 1 | reached_patient: "unknown" was read as "no", discarding a completed booking. Only an explicit "no" means not reached — unknown is not no. |
| 2 | wait_for_result returned status: "completed" while structured_result was still null. A terminal status is not a finalised result; the client now re-reads once. |
| 3 | The reference code RCR-708A4D was read aloud as "capitalized R, capitalized C, capitalized R, dash, seven, zero, eight…" — fifteen seconds nobody could write down. Now six digits, no 0 or 1, spoken as three pairs. |
| 4 | The idempotency key was derived from identifiers that did not change when the request did. Correcting a patient's name produced "Idempotency key was reused with a different request" — and no call went out at all. The key now covers the task text. |
| 5 | The caller left "about arranging attendance after one missed session" on a voicemail. Anyone who plays that message back learns this person is a cardiac rehab patient who has been missing appointments. A voicemail is not a private channel. |
| 6 | CALL-E reports a voicemail as status: "completed", so it was logged as no_answer. Detection now reads the post-call evidence text. |
And one experiment that failed usefully. Asked to make the conversation richer, I lengthened the goal: recite the attendance history, read the barrier back to confirm understanding, ask what times would work before offering. Same patient, same day, one variable changed. The richer version spent forty seconds on "Okay… Great… No rush… I'll hold" before introducing itself, heard "the bus route changed" and replied "I'm not with the clinic's password system", and never reached an offer. The recipient hung up. The short instruction, on the same patient, booked a slot in a coherent 53-turn call.
Instructions compete for attention. Complexity belongs on our side of the call, not inside it — which is why the negotiation is an enumerated ladder the caller cannot get wrong rather than a technique it has to remember. Recorded as ADR 0003.
Accomplishments that we're proud of
- 67 tests, no credentials, no network, 0.07s. Most exist because a real call went wrong.
- The safety boundary is asserted, not documented. One test poisons the
course file with
diagnosis,clinical_note,medicationandprognosisand asserts the generated call script is byte-identical. Another asserts no script contains assessment phrasing. A third asserts a patient who mentions a symptom and accepts a slot in the same call is escalated and not booked. - The anti-scam path fired unprompted on a live call. When the recipient said "I don't remember who I'm talking about", the caller did not argue — it said to hang up, ring the clinic's published number, and quote the reference. That design exists because an earlier CALL-E test call was interrogated by a family member and had nothing to say.
- We measured the platform. 43% of a real call was dead air — 73 seconds of
- That is a number the CALL-E team can use.
- Nothing invented enters the record. A concession the clinic never authorised is discarded rather than logged; the transcript still holds whatever was actually said.
What we learned
Cheap disqualifying checks come before expensive qualifying ones. The one command that killed the original idea took a second and was scheduled third.
A fixture in a shape the service does not return is worse than no fixture, because it passes. We wrote a transcript reader from another app's test file rather than the API reference. Every test passed; every real call would have returned an empty transcript. It would have surfaced for the first time on camera.
More instruction is not more capability. The single clearest measurement of the project was that making the agent's brief richer made it worse at finishing sentences.
Read the API reference, not the ecosystem. docs.heycall-e.com existed the
whole time and nothing in either GitHub repo pointed at it prominently. Two of
the six faults trace directly to working from the integration repo instead.
What's next for Rehab Course-Adherence Caller
- Exercise the concession ladder live. Covered by fixtures and tests; the one call that would have reached it was ended by the recipient first.
- Attendance from a real practice-management export, not a JSON file. The trajectory logic does not care about the source; the ingestion is the gap.
- Thresholds from a clinician, not from me. The lapse window, the call cap, the cooldown and "three attended" are demonstration defaults and the README says so. They belong to whoever runs the clinic.
- Webhooks instead of polling, now that the terminal-status race is understood.
Built With
- call-e
- calle-ai
- pytest
- python
Log in or sign up for Devpost to join the conversation.