Inspiration## What inspired us

Warranty support in India is a maze of PDFs, unread terms, and phone trees. The Right to Repair push and the National Consumer Helpline’s intake checklist make the same point: before anyone can help you, they need product, serial, purchase date, and a clear issue — and most people do not have that ready when they finally dial.

We were inspired less by “AI that talks on the phone” and more by the boring failures of a real warranty call:

  • calling before you have the serial number, then calling again;
  • an agent agreeing to a fee you never authorised;
  • hanging up with a case number you misheard;
  • a script timing out, retrying, and opening two cases;
  • a structured JSON result that asserts a reference nobody said out loud.

CALL-E gives builders a serious phone runtime. We wanted the layer around it: something that decides whether a call should happen, what may be said, and whether the answer is believable — with ordinary code and tests, not only prompt hope.

That layer is ClaimCall. The full consumer product around it is ClaimPilot; for CALL-E, ClaimCall is the entry.

What we learned

  1. Consent has to name a plan, not a vibe. Binding approval to a content hash means “I agree to this call” cannot silently become a different script.
  2. Idempotency belongs before the network. Writing a reservation before POST /v1/calls is what makes a timeout safe: reconcile by read, never by blind redial.
  3. Structured results still need evidence. CALL-E’s JSON is useful; claiming VERIFIED still requires the case reference to appear in the transcript.
  4. Dry-run must be the real pipeline. A fixture transport that shares dispatcher, reservation, and verification code is more honest than a stub demo path judges cannot trust.
  5. Allowlists beat UI secrets. Destination hashing (and budget / live flags) matter more for a public demo than a token field that looks like consumer UX.

We also learned the product lesson: laypeople want a one-page PDF keep-sake, not a JSON export — and judges want a CLI that installs clean and refuses loudly.

How we built it

ClaimCall (Python CLI under contrib/awesome-phone-call-agents/apps/python/claimcall):

  • compile purchase + warranty evidence into an inspectable CallManifest;
  • run a deterministic safety gate (seventeen checks) before anything can dial;
  • require --confirm <hash> for live mode so consent is plan-specific;
  • dispatch through CALL-E with an Idempotency-Key and a strict result schema (claimcall.result.v1) that pins fee.accepted to false;
  • reconcile via authenticated provider reads / webhooks;
  • verify structured outcomes against the transcript (VERIFIED vs NEEDS_HUMAN).

Default is dry-run: claimcall preview and claimcall run contact nobody, yet exercise the same gates. Live is opt-in and allowlisted.

ClaimPilot (same monorepo) wraps that engine in a FastAPI + React app: invoice extraction, Strands assessment graph, human authorisation UI, and a deployed fixture demo that cannot spend CALL-E credits. For CALL-E judging, the CLI and the live-call path are the centre of gravity; the web app shows the engine is reusable.

Stack in short: Python · CALL-E API · deterministic safety + idempotency · optional Strands/Bedrock/OpenAI for the product layer · AWS Amplify/Lambda for the deployed ClaimPilot demo.

Challenges we faced

Making refusals first-class. It was tempting to demo only the happy path. The harder work was shipping fixtures that fail: injected instructions, fabricated references, mid-call fees — each exiting NEEDS_HUMAN with a clear reason.

Live demos without dialling strangers. Open sign-up on a team stack meant we could not rely on obscurity. Hashed allowlists, budgets, and plan-bound approval became the real controls; we removed a live-role token that felt like consumer UX but was really an ops secret.

Popup-blocked “Save as PDF”. Async fetch broke window.open; we switched to an in-page print iframe and rewrote the PDF as plain language so the artefact matches the product promise.

Honest model degradation. Bedrock authorization was not always available; the product still completes via deterministic fallbacks so the public demo remains meaningful without implying a model ran when it did not.

Two tracks, one engine. Keeping ClaimCall installable on its own while ClaimPilot stays a full product forced clear package boundaries and a README that tells CALL-E judges where to start.

One-line pitch

ClaimCall is an evidence-to-authorized-claim-call compiler and reconciler on CALL-E: dry-run by default, consented and allowlisted when live, and unwilling to report a case number the transcript never said.

Share this project:

Updates

Submission history