Inspiration

About 100 million Americans carry medical debt, and a large share of nonprofit hospitals still bill patients who are likely eligible for free or discounted care under their own charity-care policies. The people who fix this today are nonprofit patient advocates and hospital social workers, and every case means reading an itemized bill line by line against coding rules, price files, the No Surprises Act and a hospital's financial-assistance policy. That is exactly the routine, repetitive, judgment-heavy work that burns out the few people doing it, so we built an agent that does the reading and the paperwork while the advocate keeps the final say.

What it does

Countercharge takes an itemized hospital bill (plus the insurer's explanation of benefits and the household's size and income) and audits every line against CMS's public rulebooks: NCCI procedure-to-procedure edits, Medically Unlikely Edits, hospital price-transparency files, No Surprises Act protections and IRS 501(r) charity-care rules. Each finding comes with a citation to the exact dataset, version and record, and a dollar amount. The agent explains the findings in plain English, drafts a dispute letter or a charity-care application that may only cite proven findings, pauses for a human to approve, sends it, and schedules a follow-up. Advocates get a queue of cases and an approvals inbox; the Trust Center shows the live policies and every refusal.

How we built it

The core is a deterministic audit engine in Python built on 3.2 million real NCCI edits, the current MUE tables, the 2026 physician fee schedule, HCPCS Level II, the 2026 federal poverty guidelines and a real hospital price-transparency file; it never calls a language model. The case agent is built with Strands Agents: an orchestrator with specialist sub-agents for bill extraction, charity-care research and letter writing, hooks that capture policy denials, and Strands interrupts that stop the agent before any outside action until a person approves. Every tool lives behind Amazon Bedrock AgentCore Gateway (MCP, Cognito JWT) with an AgentCore Policy engine in enforce mode, so Cedar rules decide what each role may do: no letter without cited findings, no recipient outside the hospital's billing domain, no escalation by patients, amounts above a cap only for advocates. Behind the gateway, Lambda tools verify KMS-signed findings and an approval token bound to the exact letter, recipient and amount, so a letter changed after approval is rejected. Case state lives in DynamoDB, sessions in S3, letters go out through Amazon SES, and follow-ups run on EventBridge Scheduler. The model provider is pluggable (Bedrock, OpenRouter or any OpenAI-compatible endpoint).

Proof

Every package is tested: engine 121, core 82, tools 71, policies 90, agent 64, evals 42. We pre-registered an answer key (hash committed before the first run) for 45 synthetic bills rendered as realistic PDFs: every seeded error was found with the exact dollar amount, and the 8 clean bills produced zero disputes. A second evaluation reads the rendered bill images with the vision model: every line item was extracted exactly, and 86.7% of cases reproduce the exact disputed amount end to end from the photo, with zero false disputes; the misses come from two issues we found and are fixing (minus-signed adjustments and matching the hospital from its printed name). A live AgentCore Policy engine returns Tool Execution Denied for uncited, unapproved or misdirected letters, and a recorded live agent run shows audit → grounded draft → approval interrupt → send.

Challenges we ran into

Our AWS account has AgentCore Runtime, Memory and Browser quotas at zero and Bedrock model access disabled, so the agent container keeps the exact AgentCore Runtime contract but runs alongside it, while Gateway, Policy and Identity are real AgentCore resources. Getting Cedar to express "only cited, approved, correctly addressed letters" taught us to split the job: Cedar enforces the shape of every call, and the Lambda verifies the cryptography Cedar can't see. MUE adjudication indicators and overlapping findings also forced us to rethink how to count disputed dollars without ever overclaiming.

Accomplishments that we're proud of

The agent cannot hallucinate a dispute: amounts come from a deterministic engine, the letter is checked against the cited findings before it can be drafted, and three independent gates must all agree before anything leaves the building. The graded evaluation is pre-registered, and the honest limits are published alongside it.

What we learned

For people handling other people's money and medical records, trust comes from refusals you can see: the most persuasive screen is the one where the agent is told no and explains why.

What's next for Countercharge

Deploy the agent on AgentCore Runtime with Memory and Browser once our quotas are raised, add voice intake so patients can simply describe their bill, read bills and hospital replies straight from Gmail through AgentCore Identity, pilot with a patient-advocacy nonprofit, and add inpatient DRG checks and more hospitals' charity-care policies.

Built With

Share this project:

Updates

Submission history