Inspiration

Most consumer-facing phone systems are exactly as frustrating as everyone assumes — anyone who's sat on hold with a healthcare provider's IVR knows the feeling of a system built to route you around a problem rather than actually solve it. That made voice-calling agents a genuinely interesting space to build in, but a lot of the conversational-AI voice demos we'd tried before CALL-E felt like impressive audio technology without a real task behind it — convincing speech, nothing actually accomplished at the end of the call. CALL-E's model — plan a call, dial out, hold a real conversation, come back with a structured result — read like the first one actually built to get something done, not just sound good doing it. So the goal became testing that claim against a real, bounded task rather than a synthetic demo: an eBay Money Back Guarantee dispute, where re-explaining the same facts, knowing which policy actually applies, and knowing how far to push before deferring to eBay's own review process is a genuine, if minor, everyday hassle. It's also a good stress test for a harder question that has nothing to do with audio quality: how do you let an agent act on your behalf on a live call without ever letting it agree to more than you actually authorized?

What it does

  1. Intake — a buyer describes the dispute: what the listing claimed, what actually arrived, the order details, and exactly which resolutions they're willing to accept (a fixed set of options, never free text).
  2. Policy retrieval — the case facts are matched against a small curated corpus of real eBay Money Back Guarantee policy documents.
  3. Recommendation synthesis — Claude assesses how specific the listing's claims were, calibrates confidence accordingly (a bare "Used" label is a weaker case than a specific, contradicted claim like "plays perfectly"), and drafts a call script with verbatim policy citations.
  4. Human approval — the buyer reviews the recommendation, the citations, and the exact call script before anything happens. The script is the entire behavioral contract CALL-E receives — there's no mid-call intervention once it's placed, so everything has to be right here.
  5. The call — CALL-E places a real outbound call to eBay customer support, states the claim, cites the specific contradiction, and holds to the pre-authorized settlement scope.
  6. Outcome review — the transcript and a structured result (resolved / escalated / no resolution, with a specific escalation reason) come back for review.

How we built it

Next.js on the frontend, AWS Amplify Gen 2 for the backend. The system is six logical components: an intake UI, a case store, policy retrieval (the corpus is small enough that "stuff it into the prompt" beats standing up a real embeddings pipeline), a recommendation synthesizer (Claude), a human approval gate, and a CALL-E adapter plus webhook receiver. The synthesizer and the CALL-E adapter are separate Lambda functions. The webhook receiver writes directly to DynamoDB via the AWS SDK rather than through Amplify's generated Data client — a deliberate change made after Amplify's schema-level access-grant mechanism proved unreliable for that function during development.

Testing was staged rather than jumping straight to a live dispute call: a no-call dry run of the full pipeline first, then self-test calls where CALL-E called our own phone to validate the escalation logic cheaply, then a manual recon call to real eBay support (to learn the actual IVR/verification flow before spending call budget on it), then a CALL-E-placed recon call to the same number, and only then the real demo dispute call.

Challenges we ran into

  • A ~36–48 hour CALL-E platform outage blocked every call placement attempt for the better part of two days — independently corroborated by other hackathon participants on Discord with the identical failure signature, across different accounts and destination numbers. We ruled out three candidate causes in our own code before concluding it wasn't ours to fix, which cost real time to confirm.
  • A real, reproducible phone-number bug. A phone number embedded in the call script as a raw digit string got read aloud by CALL-E's voice synthesis in a garbled, non-standard grouping — and failed eBay's own account lookup on the first live attempt because of it. Reformatting it to standard dashed notation fixed it immediately; the very next call correctly navigated eBay's full IVR tree.
  • A real model self-consistency bug. A structured call result set escalation_triggered: false while simultaneously reporting an escalated outcome — an internal contradiction the schema's required/enum constraints didn't prevent. Fixed with both a schema-level hint and a defensive client-side check that never trusts either field in isolation.
  • Amplify's access-grant mechanism for the webhook Lambda's DynamoDB writes proved unreliable enough during development that we switched that one function to writing via the AWS SDK directly, rather than losing more time to it.
  • "Third-party representation" turned out to be two separate problems, not one — account access (solvable with per-user OAuth) and phone-call identity verification (which a real recon call showed eBay checks with just name, phone, and zip — no login required at all). Untangling which was which took longer than expected, but shaped the design meaningfully.

Accomplishments that we're proud of

  • Every phase of testing used real placed calls, not mocked responses — including a full run against eBay's actual live phone system.
  • All three designed escalation conditions (a representative declining to act, a representative proposing something out of the authorized scope, and a representative disputing the underlying facts) were each confirmed firing correctly in live calls.
  • Two real, non-obvious bugs were found and fixed — the kind that only show up when you actually place live calls rather than mock the integration.
  • We wrote up the honest limitations rather than a sanitized happy-path story — including the platform outage, the caller-ID mismatch, and what remains genuinely untested.

What we learned

  • In an async, task-in/structured-result-out integration with no mid-call intervention, the call script is the entire behavioral contract. There's no "we'll fix it live" — precision in what you hand the agent up front matters more here than in a typical text-based agent.
  • Don't trust a single field of a structured LLM output in isolation when other fields in the same object encode overlapping information — cross-derive defensively instead.
  • An outbound voice AI calling on someone's behalf carries a different trust burden than an inbound one. The recipient has to judge, in real time, whether an AI-voiced call with a mismatched caller ID, claiming to act on someone's behalf, is legitimate — and that's a harder problem to solve by engineering the calling system alone.
  • Platform outages happen mid-hackathon and can eat a full day of your schedule. Independently corroborating a suspected platform issue (in our case, via Discord) before assuming it's your own bug saved us from chasing ghosts in our own code.

What's next for Fair Call

  • Per-user eBay OAuth in place of the manual intake form, so each customer authorizes their own account directly instead of going through one shared account.
  • A real vector-retrieval RAG layer if the policy corpus grows beyond what comfortably fits in a single prompt.

Built With

Share this project:

Updates

Submission history