Inspiration

Did you know that an estimated 80% of all medical bills in the United States contain billing errors? When hospitals "upcode" a routine triage visit as high-complexity emergency care (CPT 99285) or "unbundle" blood panels that should be billed together under comprehensive codes, everyday patients are left holding the bag for thousands of dollars in overcharges. The system is intentionally opaque, and fighting these errors requires deep knowledge of CMS (Centers for Medicare & Medicaid Services) coding rules—knowledge the average person simply doesn't have.

We were inspired to level the playing field by building an autonomous AI agent that acts as a personal, expert medical billing advocate for patients, giving ordinary families the clinical and legal leverage to fight predatory medical debt.

What it does

MedAudit PRO is an autonomous, event-driven background agent built for the Everyday Agents track.

Instead of forcing users into another tedious chatbot conversation, MedAudit PRO operates ambiently:

  • Zero-Noise Ingestion: Patients drag and drop their medical bill, EOB (Explanation of Benefits), or 837P statement into our dashboard.
  • Autonomous Cross-Examination: The agent autonomously parses the bill, cross-references every CPT code against official CMS Medicare Physician Fee Schedules (PFS), and checks for unbundling infractions using CMS National Correct Coding Initiative (NCCI) bundling edits.
  • Automated Dispute Generation: When billing violations or excessive price markups (>2.5x Medicare baselines) are detected, the agent autonomously generates a formal, regulatory-compliant dispute letter citing federal authorities like the No Surprises Act (Public Health Service Act § 2799A-1) and the False Claims Act (31 U.S.C. §§ 3729–3733).
  • Human-in-the-Loop Approval: The patient reviews the flagged line items, inspects the calculated savings in the "Dispute Desk" Action Modal, and authorizes dispatch with a single click.

How we built it

We built MedAudit PRO using an enterprise serverless architecture on AWS centered around the Strands Agents SDK:

  • The Cognitive Agent Core (Strands Agents SDK): Built with the Strands Agents SDK (Python) and integrated with Amazon Bedrock (supporting Bedrock Mantle openai.gpt-oss-120b and Claude 3.5 Sonnet). The agent utilizes a structured Reason-Verify-Decide protocol equipped with three custom @tool functions:
    • query_policy_rules: Verifies insurance plan rules, coinsurance rates, and prior-authorization status.
    • check_unbundling: Cross-examines billed CPT codes against CMS NCCI panel guidelines to detect unbundled lab sets.
    • draft_appeal_letter: Generates formal Markdown dispute notices citing statutory federal protections.
  • Document OCR: Amazon Textract (AnalyzeDocument with Tables and Forms) extracts structured billing line items and CPT codes from messy PDF invoices.
  • Backend API: FastAPI (Python 3.12) packaged in custom Docker containers and deployed on AWS Lambda via Amazon API Gateway.
  • Database & CMS Reference: Amazon RDS (PostgreSQL) storing historical CMS Medicare fee schedule baselines and NCCI bundling edits, managed asynchronously with SQLAlchemy and Alembic.
  • Secure Storage: Direct-to-cloud PDF uploads to Amazon S3 via presigned URLs, triggering S3 event notifications.
  • Frontend Console: Built with React, Vite, TailwindCSS, and Framer Motion for a sleek, dark-mode "Dispute Desk" console deployed on Vercel.

Challenges we ran into

  • Bypassing the 29-Second API Gateway Timeout: Running Amazon Textract OCR paired with multi-step Strands Agent tool reasoning can take 30–45 seconds on complex multi-page bills. Because Amazon API Gateway enforces a strict 29-second timeout, we decoupled ingestion: the API Gateway delegates to an asynchronous S3 event dispatcher, and the backend communicates with the Strands LLM Lambda using direct Boto3 Lambda invocations (RequestResponse), completely bypassing gateway timeout limits.
  • Eliminating AI Hallucinations in Medical Coding: In medical billing, an invented CPT code or fictional pricing baseline destroys credibility. We eliminated hallucinations by strictly constraining the Strands agent prompt persona and equipping it with deterministic @tool functions that pull ground-truth data from our RDS CMS database.
  • Dual Provider Resilience: To protect patients from third-party model downtime or token rate limits, we engineered a three-tier execution gateway: Mode A (Bedrock Mantle Proxy), Mode B (Native AWS Bedrock Runtime), and Mode C (Deterministic Heuristic Fallback Gate).

Accomplishments that we're proud of

  • 100% Benchmark Accuracy: We engineered an automated evaluation suite (agent/evaluation/eval_accuracy.py) testing the agent against clean bills, emergency upcoded visits (CPT 99285), and unbundled metabolic panels. The Strands agent achieved 100% precision and recall (3/3 test cases passed) with zero false positives.
  • Seamless Zero-Noise UX: Created a true "background agent" experience—patients don't have to prompt the AI; the agent does the hard clinical and legal work silently and presents the finished dispute ready for signature.
  • Full Monorepo Architecture: Seamlessly organized the frontend, backend, agent engine, and serverless infrastructure using Git submodules in a clean, reproducible open-source monorepo.

What we learned

Building with the Strands Agents SDK showed us the profound difference between a conversational chatbot and an autonomous agent. By giving the LLM deterministic tools and grounding it in official government datasets, AI can navigate complex, asymmetric bureaucracies on behalf of humans. We also mastered chaining AWS serverless primitives (S3 events, Lambda containers, Bedrock, and Textract) into a robust production pipeline.

What's next for MedAudit PRO

  • EHR Portal Integration: Direct OAuth integration with patient health portals (Epic MyChart, Cerner) to automatically audit new medical statements as soon as they are issued.
  • Direct Dispute Auto-Dispatch: Integrating Amazon SES to electronically dispatch approved dispute letters directly to hospital billing compliance departments and state insurance commissioners.
  • Expanded NCCI Procedure Tables: Scaling our local CMS database from urgent care and laboratory panels to comprehensive surgical and inpatient diagnostic codes.

Built With

Share this project:

Updates

Submission history