Inspiration
A hospital bill arrives as a sheet of codes and dollar amounts. 41% of US adults carry debt from medical or dental bills, and 51% of those say cost stopped them getting a test or treatment a doctor recommended (KFF Health Care Debt Survey, 2022).
The document that could help already exists. 45 CFR 180.50 requires every US hospital to publish a machine-readable file of its own prices. The files are public, free, and enormous. Norman Regional Health System's is 39,141,747 bytes and 151,439 lines. OU Health publishes a wide CSV with 467 columns through a vendor endpoint. Almost nobody opens them. A hundred and fifty thousand rows of billing codes is not something a person reads. That is a job for an agent.
What it does
You photograph a hospital bill. Fairbill reads it, identifies the hospital, downloads that hospital's published standard-charges file live, and puts the hospital's posted price next to every line you were charged.
It checks five things: the same code billed twice, a lab panel split into its components, a visit level billed higher than the notes support, a charge above the posted cash price, and an out-of-network charge that federal law protects you from.
When it finds something, it shows the row. Bill line 5 says $449.00. The hospital's file, row 14756, says the cash price is $269.40. Then it hands you one card: send the dispute letter, request the itemized bill first, or pay as billed. The letter is already drafted, cited to the regulation, and addressed.
If the bill is clean, it says so. That case is in the demo on purpose. An auditor that always finds something is not an auditor.
Fairbill never moves money. There are no payment fields anywhere in the product.
How we built it
Three specialists run in parallel inside a Strands Graph: a line matcher, a coding specialist, and a rights specialist. Each proposes claims using structured output. A claim names a kind and the relevant billing codes. It never carries a price or a citation. A deterministic Python validator then re-derives every finding from the hospital's own file. If it cannot rebuild the claim, the claim is dropped. That is why the fallback model cannot lower the evidence bar: the bar is in the code, not the model.
A context node fans out to the three specialists, which converge on the validator gated by all_dependencies_complete. Each specialist subclasses MultiAgentBase and never raises: a failed specialist retries on the fallback model and returns an empty claim set.
The reader uses OpenCV to find the page quad and warp it flat before the model sees anything. Two passes follow: a verbatim transcription, then a structured-output pass into a Bill schema. Deskew took the reader from 3 of 6 to 6 of 6.
A Strands InterventionHandler with on_error = "deny" guards every tool call. It denies anything that touches payment verbs, card-shaped digits, or routing numbers. A decoy pay_bill tool exists to be denied on camera. Its execution counter has never left zero.
Primary model is Claude Sonnet 4.6 on Amazon Bedrock (us.anthropic.claude-sonnet-4-6). Fallback is Claude Haiku 4.5, same region. Both pass through the same validator.
The engine runs on an Amazon Bedrock AgentCore Runtime (direct-code zip, Python 3.13, ARM64) behind a Lambda Function URL in streaming mode. One URL, one QR code.
Challenges we ran into
A few degrees of rotation breaks a statement. On a wide bill table, a three-degree tilt shifts the money columns by a full row height and the model attaches the wrong amount to the wrong line. The fix was geometry (OpenCV deskew), not a better prompt.
Models will cite anything. A review harness fed the first version forged claims and it accepted 7 of 12: cross-date duplicates, above-cash flags on insured lines, two-component unbundling, negative amounts. Each is now a rule in the validator.
Hospital files are not one format. Norman publishes a tall CSV. OU publishes a wide CSV with 467 columns. The parser is written against the CMS data dictionary rather than either file.
Accomplishments that we're proud of
The reader reads 6 of 6 demo bills with every code and every amount exact. The audit finds 5 of 5 planted errors and raises 0 false flags, including 0 on the clean bill, from both the printed bill and the phone photo (10 of 10 overall). Every drafted letter scores 1.00 from the Claude judge. The decoy payment tool has never executed.
What we learned
Parallel is not the interesting part of a multi-agent system. Separation of powers is. Letting models propose and letting code decide gave a harder guarantee than any amount of prompting, and it made the fallback model safe to use.
The reader is the whole product. If a digit is wrong, every downstream step is confidently wrong. Getting the image geometry right mattered more than anything we did to the prompt.
Refusing to find something is a feature. The clean bill does more for trust than any of the five bills with errors.
What's next for Fairbill
An MCP server exposing audit_bill so other agents can call the engine. AgentCore Memory for standing rules, such as "always request the itemized bill first." Background re-auditing when the itemized bill arrives. Coverage beyond the five error types and the hard-coded lab panels, and a real user test with people who have a bill they cannot check.
Built With
- amazon-bedrock
- python
- strands-agents
Log in or sign up for Devpost to join the conversation.