Project title: BillRx

Tagline: An evidence-first agent that helps people investigate confusing medical bills.

Inspiration:

Medical bills are some of the most confusing documents ordinary people ever receive. Itemized statements, explanation of benefits forms, procedure codes, and remark codes: it is a paperwork maze, and most people just pay whatever the statement demands. I wanted to build an agent that reads a bill the way a forensic accountant would, line by line, and tell you exactly which charges are worth questioning and what to ask about them.

What it does:

BillRx takes an itemized hospital bill, and the insurer's explanation of benefits audits every line and surfaces charges worth questioning. Each finding comes with exact evidence citations and a specific question to ask the billing office. It then drafts a formal appeal letter and a phone script and pauses at a human approval gate. Nothing is finalized until a person reviews and approves it.

The demo runs on a synthetic scenario with four candidate findings: a duplicate lab charge, a component panel billed alongside its comprehensive panel, a statement balance that far exceeds what the EOB says is owed, and one clean venipuncture line that the reviewer honestly rejects for lack of evidence. The claim delta in the fictional example is a $1,842 difference worth investigating.

How we built it:

BillRx is built with the Strands Agents SDK and follows one architectural rule:

The model never calculates the money. An extractor agent structures the bill and EOB. A deterministic checker written in plain Python runs every number with exact decimal math: it totals the lines, hunts for duplicates, checks billing code relationships against a reference table, and compares the statement balance with the EOB patient responsibility. The candidate findings it produces are hypotheses, not verdicts. A reviewer agent then challenges each hypothesis and may only judge after calling lookup tools that return facts. A drafter writes the appeal and phone script from templates plus the reviewer's reasons. Every stage emits events to an audit trace shown on the Investigation screen.

Two modes share one contract. Mock mode uses a deterministic scripted reviewer. So the demo is fully reproducible. Live mode runs the reviewer as a real Strands agent on Amazon Bedrock (verified with Nova Lite in us-east-1), with a deterministic fallback if the live reviewer ever fails to call its tools.

Challenges I ran into:

The hardest problem was verdict calibration in the live agent. A small model asked to KEEP or REJECT findings, kept hedging toward rejection even when it was its own reasons, and described a textbook question worth asking. Four rounds of redesign moved the behavior but never made it match the deterministic contract. So we made an honest engineering call: the recorded demo uses the deterministic mock mode, and the live Bedrock path ships as real, working code with its behavior documented. Reliability beat cleverness.

What's next?

With more time I would add OCR for scanned bills, a larger reference table of billing edits, and PDF export of the dispute packet. The architecture already supports all three: the deterministic checker is the seam where new rules plug in, and the approval gate is where new outputs wait for a human.

Disclaimer:

BillRx is a research prototype using synthetic billing data. It does not determine whether a real medical bill is incorrect and does not provide medical insurance or legal advice.

Built with Strands Agents SDK, Amazon Bedrock, Python, Streamlit, and pytest.

Built by Akram Mohammed.

Built With

  • amazon-bedrock
  • pytest
  • python
  • strands-agents-sdk
  • streamlit
Share this project:

Updates

Submission history