Inspiration
Chargebacks are decided on evidence, and the evidence that wins is the customer's own words: "yes, I got it," "yes, that's my charge." Most small merchants never have that. The order email goes unanswered, nobody has time to pick up the phone, and they concede. Calling the customer is the obvious human move, and the one nobody makes. I wanted an agent that makes that call, and never files an answer the customer didn't give.
What it does
Rebuttal handles a card dispute end to end. Its coordinator, GPT-6 Astra on the OpenAI Agents SDK, gathers the dispute from Stripe, the order record and the customer's email thread. If the thread is silent, it places one CALL-E call with a fixed script that opens by saying it's an automated assistant calling for the merchant, and asks two questions: did you receive the order, and do you recognise the charge.
Then it does the part a completed call doesn't guarantee. CALL-E's structured result is treated as a claim and cross-examined against the transcript:
- a "yes" is used only if a customer turn answering that question actually says yes
- the answers count only if the disclosure was spoken, the call completed, and CALL-E's completion confidence is at least 0.8
- the caller must never have asked for payment details
- a "no" stops the filing even when it isn't grounded: a yes needs the customer's words, a no only needs to be possible
A call that survives becomes a document: the masked number, every turn with its offset, and every check with the quote that passed it. Rebuttal uploads it to Stripe as dispute evidence, puts the quotes into the rebuttal, asks a human to approve in Slack, reads Stripe's Visa Compelling Evidence 3.0 validator, and files once.
Calling is where an agent can do the most harm, and a placed call can't be recalled. So six calling rules run before CALL-E ever sees the request. The number must be the customer's number on record, never one the model supplies. A live call needs per-run operator intent and an allowlisted destination. The recipient's local time must be between 08:00 and 21:00. The script must be byte-identical to the template. And a dispute is called at most once, backed by a CALL-E idempotency key.
How I built it
- CALL-E Python SDK (
calle-ai), called at runtime.calls.createwith a strictresult_schema(received, recognises_charge, purchaser, declined_to_talk),metadataand an idempotency key derived from the dispute.calls.list_eventsstreamed into Slack while the phone rings.calls.getfortranscript_turns,structured_resultandcompletion_confidence. - GPT-6 Astra on the OpenAI Agents SDK, coordinating eight tools across Stripe, Google Sheets, Gmail, CALL-E, Slack and Photon. The verdict comes from a policy table, not from the model.
- One write-gate with 13 forbidden effects, six of them about calling, enforced in code and traced as attempted → blocked → reason.
- Grounding in plain Python (standard library, plus reportlab for the PDF), so the call module lifts into any agent unchanged.
- Twins of all six apps, including a CALL-E twin that mirrors the live payloads and idempotency, so the real coordinator runs a 28-scenario suite without placing a single call.
- A silent-failure detector that names Hallucination, Instruction Violation, Skipped Work and more on every run.
- A Next.js site on Vercel that replays saved CALL-E calls turn by turn, with a FastAPI read API on Render.
Challenges I ran into
- My first live call looked perfect. It completed at 0.95 confidence with yes and yes, and the customer really did say yes. But when I read the transcript, the caller never said it was automated. So I built the cross-examination, rewrote the script and made disclosure a condition of use. That call now lives in my tests as the one Rebuttal refuses to file.
create_and_waitblocks for the whole call with no progress, so I switched to create plus event polling to show the call as it happens.- Speech recognition is noisy ("HI received order."). Grounding matches meaning with simple patterns and treats anything unclear as unknown, which changes nothing.
- A submitted call can't be cancelled through the public API, so every calling rule has to run before
calls.create, not after. - Stripe's disputes API submits by default, so Rebuttal always stages with
submit=falseand reads the validator first.
Accomplishments that I'm proud of
- An answer from a phone call has to be backed by the customer's own words before it ever reaches a bank.
- Calling rules that stop a bad call before it exists: a call at 23:00 in New York, to an unauthorised number, or with an edited script never reaches CALL-E.
- 52 unit tests, plus 28 seeded scenarios run by the real coordinator, 9 of them calls. In one, CALL-E reports yes and yes while the customer said "Sorry, who is this?" Nothing is accepted, and the detector names it a hallucination.
- A no-key path anyone can run in a minute:
python -m rebuttal.confirm. - I contributed the call to CALL-E's awesome-phone-call-agents as a standalone app with 75 tests and a dispute-evidence skill.
What I learned
A completed call is not a trustworthy answer. The most dangerous failure isn't the call that fails; it's the confident one that's true but was obtained in a way you can't use. And a "yes" and a "no" deserve different burdens of proof.
What's next
- Calls to carriers and couriers for proof of delivery
- Consent-based recording retention, and disclosure and recording rules per jurisdiction
- Grounding beyond English
- A pilot with small merchants who concede disputes today
Log in or sign up for Devpost to join the conversation.