Inspiration

In Bangladesh, and across South Asia, most online orders are paid in cash when the courier hands them over. Nothing is charged when the order is placed. That means placing an order costs the customer nothing, and walking away from it costs them nothing either.

Shops live with the result. A large share of cash-on-delivery orders are refused at the door: wrong address, changed mind, or an order somebody never really meant to place. The shop pays the courier to carry it out, pays again to carry it back, and receives nothing.

This is not a small leak. Pathao, one of the largest couriers in Bangladesh, puts most businesses at a 75 to 85 percent delivery success rate, which means fifteen to twenty-five of every hundred parcels come back. It prices each failed delivery at Tk 195, outbound plus return, and its own recommended fix is pre-delivery communication. That is exactly this call.

So every small shop does the same thing. Somebody sits down with a phone and calls every order before dispatch. It is an hour a day, it happens late because nobody enjoys it, and on busy days orders go out unconfirmed. That call is four questions and a fixed script. It should not need a person.

What it does

COD Confirm reads the pending order book and, for each order it decides is worth calling, rings the customer through CALL-E. On the call it reads back the items and the amount due to the courier, asks whether the customer still wants the order, reads the delivery address back and takes a correction if it is wrong, and asks what time of day suits them. Then it writes a decision against the order: confirmed, cancelled by the customer, still pending, or needs a human.

The part that makes it different is what happens before the call. Calling the whole order book is the wrong default once calls cost money. A call is an investment against one specific loss, and that loss is not the order value, because refused goods come back. It is the freight, both ways. That is also how Pathao prices a failed delivery.

So each order is priced first. Risk comes from what the shop already knows: has this customer taken delivery before, have they refused before, how much is waiting at the door, how far does the parcel travel. Orders where the call costs more than the risk it removes are left alone, and the run says which ones and why.

id      total  ship   risk  saving     net  call?
1044     9800   150    95%   171.0   165.0  yes    refused twice before, out of city
1042     6200   220    25%    65.0    59.0  yes    new customer, bulky parcel
1045     4750   190    28%    62.7    56.7  yes
1041     3450    80     8%     7.4     1.4  yes    loyal, but freight is real
1043     1150    60     8%     5.5    -0.5  NO     the call costs more than it saves

Order 1043 is a loyal customer, a light parcel and a small sum. Calling her is a small, certain loss. The sweep leaves her alone and says so.

It also runs against a live WooCommerce store, not only the demo book. One setting points it at the shop. It reads the cash-on-delivery orders waiting for dispatch, prices each call from that customer's real history in the same store, deliveries taken and deliveries refused, and on a live run writes the decision back into the order as a private note. It never changes an order's status or its address; that stays the shop's decision. I checked it end to end against a real WooCommerce store holding fictional orders. CALL-E has merged that connector too, as a second pull request, #465, together with a quickstart guide in #464.

How I built it

Python, the CALL-E Python SDK, and a strict result schema.

The design rests on one decision: the agent is never asked what happened, only which of a small set of things happened. confirmed is an enum of yes, no and unclear. An agent that returns a paragraph puts a human back in the loop reading it. An agent that returns no with a reason can cancel an order on its own.

On top of that sits an evidence gate. A yes is accepted only with confirmation_quote, the customer's own words, and the brief tells the agent not to fill that in from a hum, a pause, or a yes it offered them itself. A confirmation nobody actually spoke goes to a person, because dispatching on it is how a shop ends up arguing at somebody's door.

unclear is a first-class answer and it never retries. A second call rarely produces a clearer one.

The decision table is total: every combination of answers maps to exactly one status, so no reply can leave an order in limbo.

Nothing dials unless it is told to, and then only numbers it was told about. The sweep runs dry by default, and a live run rings only the numbers on an allowlist for that run. There is deliberately no setting that means "call whatever the order book says", because an order book is data, and data can be wrong, stale, or somebody else's.

The rule I care most about is what counts as a call that did not happen. Only a request CALL-E refused outright qualifies. A timeout, a dropped connection, a call that ended as anything other than completed, or a completed call with no result are all treated as ambiguous, because the phone may have rung. That order goes to a person, the sweep stops, and nothing redials it until somebody has looked. Ringing a customer twice to ask the same question is the one failure this tool must not cause.

An answer can only confirm or cancel an order when CALL-E reports the number it actually reached and that number is the order's own. Everything that comes back as free text, the transcript, the quote, a corrected address, is cleaned before it is logged or stored.

The merged app has 95 tests covering the decision table, the pricing model, destination handling, those call-safety rules and the WooCommerce connector, which is tested against a fake store that records every write, including one that drops the connection halfway through. None of them need an API key, a store or a phone line.

Every outcome has been produced by a real call. The order in it comes from the fictional demo book and the call went to my own handset, so it contains no real customer's details. This is the transcript CALL-E returned:

bot   Hi Farhana, this is Nokshi Home calling to confirm your
      cash-on-delivery order before dispatch. Can you hear me?
you   Hi, how are you? Yes.
bot   Your order is order number one, zero, four, five: Ceramic dinner set,
      sixteen piece, with four thousand seven hundred fifty BDT due to the
      courier, for House fifteen, Zindabazar, Sylhet three, one, zero, zero;
      do you still want this order?
you   Yes, I still want it.
bot   Is the delivery address on file correct?
you   Yes, that is correct.
bot   What time of day suits you for delivery?
you   after six in the evening
bot   Thank you, bye.

-> confirmed: Confirmed, address unchanged. Said "Yes, I still want it."

That result was recorded before the safety review. The current version treats an answer from a redirected number as advisory, so the same call today would be kept and read, and a person would confirm the order rather than the agent.

An earlier call on the same order came back needs-human, because the agent could not verify a corrected address. That is the system working: it would rather hand a nine thousand taka parcel to a person than guess where it goes.

Challenges I ran into

The pricing model took two attempts. The first version priced the loss as the order value, which made every order look identical, because the freight is the same whatever is inside the box. Once the loss became freight both ways weighted by per-order refusal risk, the model started producing different answers for different orders, which is the whole point.

Three defects only appeared once real calls were placed, and none of them could have been caught by a test. place_call caught a CalleError that does not exist in the SDK. The idempotency key was the order id and its attempt count, which repeats whenever the order book is reset, and CALL-E answers 201 Created with the original call, so a replay is indistinguishable from a fresh dial and two sweeps silently returned the same stale result. And the agent was quietly dropping the address question while still returning a confident structured result, reporting address_correct as unclear because it had never asked.

The last one was the most interesting, because the fix was not code. The brief now names three mandatory questions and states in as many words that unclear means asked-and-unanswered, never not-asked.

Then the code went through five rounds of review before CALL-E merged it, and the reviewer found four gaps I had missed. There was a setting that let a run dial the whole order book. A call that ended as failed was being treated as one that never happened, when all failed really tells you is that nobody knows. A call that reported no destination was being accepted as having reached the right person, because silence looked like agreement. And a customer's quote and corrected address were being stored exactly as the model returned them. Every one of those is fixed and has tests now. None of them would have shown up in a demo.

Connecting a live store raised the same question in a new place. My first version wrote every result back at the end of the sweep. If the store stopped answering halfway, every order already called would look uncalled, and the next sweep would ring those customers again. Results are now written the moment each call ends, and a failed write stops the sweep and names the order to record by hand. A test of a store that drops the connection mid-write then caught a network error my code was not catching, which would have crashed the sweep after a phone had already rung. That one is fixed too, and the test fails on the old code. CALL-E's maintainer added one more safeguard before merging the connector, one I had missed: it no longer follows redirects, so the store's login can never be forwarded to another address.

Accomplishments that I am proud of

That the run tells you which orders it deliberately did not call, and what each of those calls would have cost against what it would have saved. Most automation is judged on how much it does. This one is worth more for the calls it talks you out of.

That it was merged into CALL-E's own repository of phone agents after five rounds of safety review, and that the WooCommerce connector and a quickstart guide were merged after it.

And that a shop can point it at the store it already runs, with one setting, and have the risk come from its own customers rather than a national average.

What I learned

That the interesting decision was not the call. Any voice API can make a call. The decision worth automating was which calls not to make, and that turned out to be a small amount of arithmetic over data the shop already has.

That a voice agent needs its instructions written the way you would write a checklist for a person who is going to be interrupted, not the way you would write a specification.

And that "the call failed" and "the call did not happen" are different sentences. Treating them as the same is how a customer gets rung twice.

What's next for COD Confirm

The per-customer signals already come from the store. Next is learning the shop-wide refusal rate from its own history too, instead of starting from a national average. After that, calls in Bangla, which is how most of these customers would rather be spoken to.

Built With

  • calle
  • calle-ai
  • pytest
  • python
  • rest-api
  • woocommerce
Share this project:

Updates

Submission history