Inspiration
Almost everyone I know has a denial letter in a drawer somewhere. A hospital episode, several claims across several providers, one or two of them refused — and then nothing happens. Not because anyone agreed with the insurer. Because the letter arrived in the worst possible week, nobody knew whether it was worth contesting, and by the time there was room to think about it the thing had gone quiet.
I watched that happen more than once, and it was always the same shape: the person had a case and no capacity to argue it, and no idea a clock was running.
So I went and read the numbers, and they say the drawer is the normal outcome. KFF's 2024 marketplace analysis reports that insurers received about 496 million claims, that roughly 85 million in-network claims were denied, and that fewer than 1% of denied claims were appealed. Of the appeals that were filed, insurers upheld 66% — so about a third were overturned.
A third of the people who push back, win. Almost nobody pushes back.
Then I read the reason breakdown, and it decided what to build. Of those in-network denials, 25% were administrative and only 5% were a lack of medical necessity. The picture most of us carry is a doctor at the insurer refusing treatment. The data says the ordinary denial is a form problem — frequently the provider's own billing error, and not a coverage decision at all.
People do not fail to appeal because they accept the decision. They fail because the appeal is unpaid administrative work, done while ill or caring for someone who is, against a deadline nobody told them about, in a format nobody explained, using documents held by three separate organisations.
And it is a clock problem. You have 180 days to file an internal appeal. The plan owes you an answer in 30 or 60 days. After a final internal denial you have four months to request an external review. Every one of those runs whether or not anybody is watching, and missing the 180 is the only failure in this domain where the merits stop mattering.
There are three good products in this space already — Counterforce Health, Fight Health Insurance, Claimable — and all three are appeal generators. You notice the denial, you go to the site, you upload the letter, and it writes you an appeal. Searching the two free products' public pages for deadline, 180 day, track, remind, monitor and ERISA returns nothing. A generator only helps the person who already knew they had been denied, already knew an appeal was possible, and already knew they were inside the window. Fewer than 1% appeal. The people that misses are not people who wanted a better letter.
What it does
Recourse closes that gap. It runs two loops over a household's case file, and the person only hears from it when there is a real decision to make.
WATCH triages what arrived. It parses an Explanation of Benefits into claim lines, classifies each denied line from the X12 CARC/RARC codes the denial itself carries, computes every deadline from the notice date in that plan's own terms, and decides the route: draft it, ask for one missing document, or hand it to a person.
That classification is the whole product, and it is a lookup rather than an interpretation. CARC 16 — "claim/service lacks information or has submission/billing error(s)" — is paperwork, and Recourse can prepare that appeal on its own. CARC 50 — "not deemed a medical necessity by the payer" — is the treating clinician's argument and nobody else's, and Recourse hands it straight back. No combination of administrative documents makes a medical-necessity denial yours.
It fails closed. A code it recognises, arriving beside one it cannot read, classifies unassessable rather than clerical, because the unreadable code might be the clinical half of the denial. Since 36% of real denials are coded "Other" with no reason listed, that is the common case rather than the edge.
PRESS runs the clock on a schedule against every open appeal. It counts your deadline and, separately, the one running against the insurer — a plan past its own 30 or 60 days has failed to give you the review it owes, and on an ERISA plan that is a door to the next stage without waiting. The console renders your clock and their clock in different colours at every risk level, because they are opposite facts and confusing them is the whole point of tracking them.
What it produces is a packet rather than a letter: the argument, the codes contested, the plan language relied on, and — beneath every sentence — the document in the case file that the sentence rests on. A letter is only worth what it can point at.
Three rules ignore the autonomy dial, and the third is the odd one out.
- It never asserts a clinical fact the record does not support.
CitationGuardreturnsDeny, notConfirm. There is nothing for a human to approve about a sentence nothing supports, and offering it would invite somebody to wave through a false statement on an insurance document made in a patient's name. - It never files. No tool in this system submits an appeal to anyone, and a test fails if one appears. Filing is a legal act taken in the claimant's name. It prepares, and a person sends.
- It never lets a deadline pass in silence. A duty, not a prohibition. Every other rule waits to be provoked by a tool call; this one fires because the agent is about to do nothing. A background agent that only spoke when provoked would go quiet through six months of a 180-day window and be right about everything except the thing that mattered.
How I built it
Strands Agents on Amazon Bedrock with Claude Sonnet, deployed on App Runner as a single container serving the agent, the API and the console.
The model reasons and never does the arithmetic. The clocks are pure functions over a notice date, a plan's rules and a stage. No model output can move a date, and no model is asked to compute one.
The rulebook and the code table are data, versioned, with each figure carrying the .gov or X12 URL it was read from — so the work can be checked without reading any Python. Two plan types ship: ACA marketplace and ERISA group health, both read from 29 CFR 2560.503-1. They share the 180 days. Where they part is what follows: an ERISA plan may require two mandatory internal levels, which splits the clock — 30 days, then 30 more, and the second only starts when the first is answered. A person tracking that by hand is tracking a clock that pauses and restarts.
A plan type with no clock table is refused, never defaulted. ERISA non-health plans allow 60 days to appeal, not 180 — §(h)(2)(i). Falling back to the health figure would tell somebody they had four months they do not have, so the loader raises and says why.
Two Strands interventions, in a deliberate order. CitationGuard runs first, because its Deny short-circuits and an unsupported clinical claim must never reach a policy that would weigh it on urgency and find nothing wrong. The autonomy policy runs second and returns Proceed or Confirm; a Confirm breaks the run out of its loop and snapshots its state, so a paused run resumes later from a structured human decision.
An MCP server sits over the same engine — read-only, and a test asserts that no served tool name contains a mutating verb. Another agent can classify a denial and price a deadline; it cannot change anyone's case.
Challenges I ran into
Everything real here I found by using the thing. Every one of them passed a full test suite first.
The insurer's clock ran from the wrong date. deadlines_for counted the plan's deadline to answer from its own denial notice rather than from the day the appeal was posted. Against the seeded world that reported Harbor Blue as ninety days late answering an appeal nobody had ever sent them — a false accusation the deadline screen would have shown as fact, on camera. theirs now runs from filed_on, which is None until a person sends the letter, so an unfiled appeal has no such clock at all and no call site has to remember to gate on it.
An appeal was "a stage plus some fields". Advancing to ERISA level two blanked level one's outcome, answer date and letter — the exact record that level two exists to contest. Worse, it blanked them by listing the fields to clear, so a field added later would have survived by being forgotten. An appeal is now the stack of levels it has passed through, advancing is an append, and history is intrinsic rather than a shadow copy the store kept beside it.
is_denied meant "the plan paid nothing", which hid every partial denial. The plan pays the visit and refuses a modifier, or cuts the units billed — those are just as appealable, and nothing in the system could see them. It now asks whether any non-cost-share code is attached, so a line paid in part with a code against it is a denial, and is_partial_denial names that case. Codes 1, 2 and 3 are deductible, coinsurance and copay; they sit on almost every paid line, and counting them would have buried the real denials in noise.
One appeal on an out-of-scope plan type turned the whole deadline screen into a 500. policy.sweep.sweep called the clocks directly and caught nothing, so a single appeal on a plan with no clock table raised straight through it. Sweep is GET /api/deadlines and the PRESS loop's sweep_deadlines, so one unreadable plan took every other appeal down with it, including the one at day 150 of 180. A household with a disability claim alongside its health ones is entirely ordinary. That is silence as a window closes — the exact failure rule 3 exists to prevent — arriving through the door marked "refuse to guess". Two sibling paths asking the same question both handled it, which is what made it invisible: /api/appeals degraded to "unassessable" on the same case file while /api/deadlines returned 500. It now refuses per appeal, says so in words a person can act on, and carries on with the rest of the file.
A closing deadline was judged as a rule against acting. The first live run against Bedrock held the agent back from drafting a letter and gave insurer_overdue as the reason. It was drafting for the one appeal whose plan is nineteen days late answering — the single case where you most want the letter written — and being owed an answer was reported as a rule against writing it. deadline_at_risk had the same shape and the same fault. It is a duty to speak, and I had it judged as a threshold on a tool call, which inverts it: the moment a window is closing is the moment to want a draft, not to interrupt one. Both now live only in policy.sweep, which raises them on a schedule at every autonomy level whether or not the agent is doing anything. The dial governs what the agent does, and warning you is not doing anything to anybody.
And the audit trail was stamped from the wall clock while everything else runs on the frozen demo clock, so the Journal screen showed 19:59 underneath a journal entry reading 09:00, on the same page.
Accomplishments that I'm proud of
The rulebook. Two bodies of appeal law read from primary sources and encoded as data with citations attached to individual figures — and a loader that refuses a plan type it has no clock table for rather than falling back to a default that would cost somebody their appeal.
The code table failing closed. That behaviour was not in my spec and should have been: a recognised code beside an unrecognised one is unassessable, not clerical. Unknown is never benign, in either direction.
The eval suite, which was asked for one honest negative result and found a real one, in the place that mattered most — the 500 above. The eval that recorded the defect is kept as its regression test.
And the seed. The agent that built the demo world rejected three of its own drafts on legal grounds: an out-of-network balance bill the No Surprises Act forbids, an annual benefit maximum the ACA prohibits, and a plan whose deductible contradicted its own claim line. A demo world that depicted an illegal bill as routine would have argued against the whole point of the thing.
What I learned
Tests tell you the code does what you wrote. Clicking tells you whether you wrote the right thing.
Driving the console found three defects that a full green suite did not. The audit-trail clock, above. /api/timeline sending detail as free-shaped JSON rather than the string its docstring implied, which crashed React the moment a real run populated the trail — invisible on the seeded world, where the timeline is empty.
And the one that mattered: the WATCH loop never cited anything. check_evidence told the agent what an appeal of that class needs and never what the case file had, so it had no document ids to reach for and every assertion came back bare. The traceability column — which is the letter's entire credibility, and the thing I would defend hardest — read "no document needed" on every run. It now returns the documents on file with their ids, and the prompt says to cite them.
I also learned to write down what I did not fix. Two findings are recorded rather than closed: CitationGuard checks that a cited document exists, not that it supports the claim, so an assertion citing an imaging bill passes; and a recognised code arriving beside an unreadable one loses the half it could read, so a person is told "unreadable" rather than "a billing error, plus something we cannot read".
What's next for Recourse
Named gaps first, because the distance between a demo and a product is where the honesty goes.
Scanned EOBs have no text layer, so the parser returns nothing and says so. A meaningful share of what people are actually mailed is a scan. OCR is the answer, and it is not here. Related: many payers print a footnote marker per row and put the code in a legend below rather than inline, and real notices batch several claims under one notice number — this takes the first. The sample PDFs are the friendly case.
The EOB is an uploaded PDF against a seeded world. There is no live payer connection. Real interoperability exists — the CMS Interoperability and Patient Access rules put claims behind a patient-authorised FHIR API — and that is where this goes next. It is not what was built here.
Nothing is durable. The stores are in memory. Restart the process and every pending decision goes with it. That is the first thing I would fix, because a background agent that forgets what it asked you is worse than one that never asked.
Then Medicare Advantage as a third rulebook row — a data file and a citation rather than a code change — and a real messaging channel behind the escalations, so the duty to speak reaches a phone rather than a screen nobody has open.
Built With
- amazon-bedrock
- fastapi
- python
- strands-agents
Log in or sign up for Devpost to join the conversation.