Inspiration

Almost everyone has a version of the same story: a confusing hospital bill shows up, the codes make no sense, and you just... pay it. We learned that most medical bills contain errors, the average large hospital bill has hundreds of dollars in mistakes, and fewer than 1% of wrongly denied insurance claims are ever appealed — not because people don't care, but because reading a bill feels impossible and fighting one feels worse. Billions of dollars a year quietly stay with hospitals and insurers who count on us not checking.

We wanted the opposite of "another app you have to manage." The Agents for Humans theme — an agent that runs in the background and only surfaces when there's a real decision — fit the problem perfectly. So we built ClaimWise: your medical-bill bodyguard.

What it does

You hand ClaimWise a medical bill — a PDF, a photo, or text — and it works in the background to find five kinds of problems:

  • Duplicate charges — the same code billed twice
  • Unbundling — two codes that can't legally be billed together
  • Preventive services wrongly charged — care that must be free
  • Overpricing — a charge far above the Medicare rate
  • Wrongful / appealable denials — including amounts that were never billable to you

For every problem it finds, it drafts a ready-to-send appeal citing the exact error and dollar amount. Then it stops and asks one question — send, edit, or skip? It never submits anything on its own, and it tracks each dispute over time instead of handing you a letter and disappearing.

Crucially, ClaimWise doesn't guess — it proves. Four of the five detectors are grounded in real, public government data, and every finding cites its source. On our sample bill it finds five issues worth $1,222.67 in seconds.

How we built it

ClaimWise is a multi-agent system built on the Strands Agents SDK:

  • An ingestion step turns a real bill into structured data — digital PDFs via pypdf, photos/scans via Tesseract OCR, and messy layouts via an LLM bill-reader agent.
  • A coordinating workflow runs the five detectors (each a Strands tool), then hands the findings to specialist agents that explain them in plain English and draft the appeal.
  • A human-in-the-loop approval gate (a real Strands interrupt) pauses before anything leaves the system.

The detectors are wired to real public datasets, each with a refresh script that pulls from the official source:

  • CMS Medicare national payment data — overpricing
  • CMS NCCI PTP edits — unbundling
  • USPSTF / ACIP / Medicare preventive-services list — free-care checks
  • X12 CARC + group codes — denials (encoding the rule that a contractual "CO" adjustment can't be billed to the patient)

We shipped it three ways: a CLI, a FastAPI web app (the live demo — upload a bill, review findings, click "Send appeal"), and an MCP server that exposes ClaimWise's abilities to any MCP host like Claude Desktop — and, through Strands' MCP client, ClaimWise's own agent can consume them. The model layer is portable (a free local model, Amazon Bedrock, or a zero-cost mock), so the whole thing builds and demos for $0 with no API key, and it's containerized with a Dockerfile for a free live URL. It's Amazon Bedrock AgentCore-ready for scheduled dispute follow-up.

Challenges we ran into

  • Getting real data, honestly. The biggest decision was refusing to ship fake lookup tables. Sourcing genuine CMS Medicare rates, NCCI edits, preventive codes, and X12 denial codes — and respecting copyright (we store code numbers and our own descriptions, never AMA CPT or X12 text) — took real work, but it's what makes ClaimWise credible.
  • Keeping the model honest. For a money-and-health tool, an AI that hallucinates a billing error is worse than useless. We made detection deterministic and auditable; the model only explains and argues findings, it never invents them.
  • OCR on messy bills. Real photos jumble table columns; we added a column-repair step and an LLM-extraction fallback so ingestion degrades gracefully.
  • Human-in-the-loop across a web request. Adapting the CLI approval gate into a stateless web flow (draft → persist → approve) so the "one decision" moment felt natural in the browser.

Accomplishments that we're proud of

  • Every finding is provable. Four detectors cite official government sources on screen — not "an AI thinks this is wrong."
  • It's a real product, not a prototype. A polished web app, real bill ingestion, dispute tracking, and an MCP integration — all working end-to-end.
  • It runs free. Anyone (including judges) can run the whole multi-agent system with no API key.
  • Genuine, on-theme Strands + MCP usage — multi-agent orchestration, a real human-in-the-loop interrupt, and MCP as both server and client.
  • It's tested — an automated suite covers all five detectors, the ingestion path, and the MCP server.

What we learned

  • Strands made "autonomous but accountable" easy — multi-agent orchestration for the work, a human-in-the-loop interrupt for the trust, and model portability so it builds free and scales on Bedrock.
  • The hardest, most valuable decisions weren't framework-specific. Grounding every finding in real public data — and keeping the model out of the "deciding what's wrong" loop — is what turns a demo into something trustworthy.
  • A lot about the U.S. medical-billing system — NCCI edits, CARC/group codes, and how contractual write-offs work — and how much of that pain is automatable.

What's next for ClaimWise

  • Live insurer/EOB monitoring and price lookups via MCP servers.
  • Autonomous appeal-deadline tracking with escalation to a second-level appeal or external review.
  • Charity-care / financial-assistance detection for nonprofit hospitals, which are legally required to offer it.
  • Broader ingestion (patient portals, emailed statements) and support beyond the U.S. system.

Built With

Share this project:

Updates

Submission history