Inspiration
Medical billing errors are incredibly common — duplicate charges, arithmetic mistakes, and bills that don't match what your insurance's Explanation of Benefits (EOB) actually says you owe. Most people never catch these because manually cross-referencing a bill against an EOB line by line is tedious, and most people don't even realize the EOB is the document that actually matters. We wanted to build an agent that does this checking automatically and silently, and only interrupts you when there's something real worth acting on — which is exactly the theme of this hackathon.
What it does
MedBill Reconciler Agent takes a patient's medical bill and their insurer's EOB (PDF or photo), extracts the line items from both, and reconciles them against each other. It catches:
- Duplicate billing — the same procedure billed twice on the same date
- Arithmetic errors — line items that don't actually sum to the bill's stated total
- Bill-vs-EOB mismatches — the provider billing more than the EOB says the patient actually owes
- Possible surprise/out-of-network billing — flagging charges that may fall under No Surprises Act protections
If nothing significant is found, the agent stays quiet — it doesn't manufacture concerns. If it does find something worth disputing, it drafts a ready-to-send appeal letter citing the specific codes, dates, and dollar amounts.
Who it's for
Anyone who gets medical bills and EOBs — which is nearly everyone eventually. No billing or insurance expertise required.
How we built it
- Amazon Textract extracts structured tables from bills/EOBs regardless of layout — different providers and insurers use wildly different formats, so a header-alias mapping layer normalizes fields like "Charge" vs. "Billed Amount" vs. "Fee" into one consistent schema.
- A deterministic rules engine (not the LLM) checks the actual arithmetic and cross-document matching — billing math is checked with math, not a language model's best guess, since accuracy matters when it involves someone's money.
- A severity-scoring tool decides whether a finding is worth surfacing to the user at all, embodying the hackathon's core theme: an agent that only interrupts for a real decision.
- Amazon Bedrock (Claude), via the Strands Agents SDK, orchestrates the five tools and handles the one task that genuinely needs language reasoning: drafting a clear, factual appeal letter once a discrepancy clears the surfacing threshold.
- A synthetic bill/EOB generator (with a PDF renderer) let us build and test the full pipeline without ever touching real patient data.
Challenges we ran into
Getting Textract's table output — which represents cells as row/column relationships to word blocks — into a clean, normalized schema that could handle real-world variation in how bills and EOBs are formatted was the trickiest part. We also deliberately kept the rules engine separate from the LLM's reasoning, so the dollar-amount findings are fully deterministic and unit-testable, rather than relying on the model to "eyeball" arithmetic correctly.
Accomplishments that we're proud of
A fully working end-to-end pipeline: real Textract extraction, a rules engine with 9 passing unit tests, severity-based surfacing logic, and genuinely well-written appeal letters drafted automatically — all wired together as Strands tools.
What we learned
How much of "agentic AI" that actually matters in a sensitive domain like billing/finance is about deciding what not to hand to the LLM. The reasoning model is best used for judgment and language tasks, while deterministic code should own anything involving exact numbers.
What's next for MedBill Reconciler Agent
The natural next step is a simple upload interface (web or mobile) instead of a command-line tool, plus deployment on Amazon Bedrock AgentCore for persistent session memory across multiple bills over time.
Built With
- amazon-bedrock
- amazon-textract
- amazon-web-services
- boto3
- claude
- pytest
- python
- reportlab
- strands-agents-sdk
Log in or sign up for Devpost to join the conversation.