Inspiration

Someone we know got a $1,840 emergency-room bill that was simply billed twice — the same visit, two lines. They almost paid it. That sent us down a rabbit hole: 85% of denied medical claims are never appealed, yet when people do appeal, about 4 in 5 get overturned. Up to 80% of bills contain errors. The system doesn't fail because patients are wrong — it fails because appealing is a second job: decode the bill and the EOB, write a formal letter with exact amounts and line references, find the right portal, then chase it for weeks. AI tools today hand you a draft and walk away. We wanted to build the thing that finishes the job — and measures itself in dollars recovered, not drafts generated.

What it does

Advocate is an AI advocate that fights US medical bills and claim denials end to end. Upload a photo or PDF — no account — and it: (1) reads the document with vision/text extraction into schema-guarded JSON; (2) runs a deterministic red-flag engine (duplicate charges, bill-vs-EOB mismatch, denial-without-reason, balance-billing / No Surprises Act signals, upcoding, missing itemization) and explains the bill in plain English; (3) asks up to 6 grounded tap-to-answer questions, one at a time; (4) writes a case brief with evidence tied to document lines; (5) drafts a formal appeal letter where every dollar amount and line reference is verified twice — a deterministic checker plus an independent semantic audit — with failures blocking approval server-side; (6) unlocks a payer-specific filing guide with print-to-PDF; (7) tracks the case with a per-case durable agent (21-day check-in alarms) until you record the outcome — won, reduced, denied with an escalation ladder, or still waiting — on a recovered-dollars result card. Nothing is ever sent anywhere without your explicit approval. Demo figures shown are synthetic cases.

How we built it

Cloudflare Workers + D1 + R2 + Durable Objects (one CaseAgent per case: alarms, reminders, purge), Hono API, React + Vite frontend, strict TypeScript, 7 D1 migrations. The pipeline runs in waitUntil so uploads return instantly while extraction → rules → findings → questions run in the background and the case page polls to results. LLMs are called over plain fetch through an OpenAI-compatible gateway (provider-swappable, no vendor SDKs; vision model verified at a fraction of a cent per case) with PDFs parsed via unpdf and scanned PDFs failing to an honest "send photos" message instead of hallucinating. Trust is structural: schema guards on every model output, deterministic rules before any LLM narrative, regex-level amount/line citation checks plus a second-model audit, a server-enforced approval state machine, signed expiring claim links instead of accounts, and a daily cron that auto-deletes cases, files, and reminders after 90 days.

Challenges we ran into

The big one was money formatting: generators repeatedly misrendered raw cent integers (301100 cents became "$301,100"), so we converted everything to dollars at the grounding boundary and re-verified with exact cent matching. Extractors also loved to "help" — inventing IDs, quoting our own finding titles as if printed in the document, asserting dates never shown — so we constrained the letter to validated fields, forced [BRACKETED] placeholders for anything unknown, and added the semantic audit as a second net. Scanned PDFs with no text layer, ambiguous EOB columns, and verifier unavailability each needed an honest failure path that never strands a case in "reading." On the product side, the hardest call was scope: voice follow-up calls were designed but cut from v1 so the letter-to-outcome loop could be genuinely complete instead of half-wired.

Accomplishments that we're proud of

A live, complete loop — upload to dollars-recovered — with no signup and a ~3-minute activation moment. A verification engine we're willing to demo adversarially: the blocked-approval state (unverifiable draft, approval locked in the API, specific reasons listed) is a first-class screen, not an error toast. Sub-cent inference cost per case against overturn rates that make every completed case return multiples of its cost. Privacy posture that holds up: expiring links, 90-day sweep across D1+R2+DO, PHI-free logs, fixed error envelopes. And honesty as a feature — synthetic demo data labeled as such, roadmap items labeled as roadmap, limitations in the README.

What we learned

AI is the easy part; trust is the product. Users won't let AI act unsupervised (and shouldn't), so we stopped trying to make the model more persuasive and started making the system more checkable — deterministic rules where facts live, models where language lives, humans where decisions live. We learned that "no result" is a result: every pipeline failure needed a designed, plain-English state, because a stuck spinner destroys more trust than a rejection. And that MVP means minimum scope, not minimum quality — cutting voice calls to polish the core loop was the best decision we made.

What's next for Advocate

  1. Voice follow-up calls on approved scripts with transcripts and logged outcomes.
  2. New dispute types on the same engine — subscriptions, deposits, tolls, chargebacks.
  3. Payer-response parsing that auto-drafts escalations (external review, state insurance complaint).
  4. Outcome-based pricing ($29–99/dispute or a share of recovered dollars), validated by the in-product willingness-to-pay survey running on every outcome card today.

Built With

  • ai-agents
  • cloudflare-d1
  • cloudflare-r2
  • cloudflare-workers
  • durable-objects
  • llms
  • typescript
Share this project:

Updates

Submission history