Inspiration

Last year insurers denied about 85 million in-network claims on HealthCare.gov. Consumers appealed 262,982 of them, under 1%, and insurers upheld 66% of the appeals they received (KFF, 2024). When the same kind of denial reaches an independent physician through California's Independent Medical Review, 72.3% were overturned in 2025.

An insurer reviewing its own denial upholds it two times in three. An independent doctor overturns it nearly three times in four. The entire gap is people not filing.

Part of why they do not file is a clock nobody tells them about. Under Cal. Health & Safety Code § 1374.30(j)(3), an enrollee "shall not be required to participate in the plan's grievance process for more than 30 days." So thirty days after you file a grievance, whether or not your plan has answered, a six month clock starts. If the plan takes four months to say no and you count six months from the no, you are already three months late. That is where the name comes from.

What it does

Day Thirty works a California health insurance denial end to end, in the background, and interrupts only for one decision.

  1. Amazon Nova reads the denial letter and pulls out the condition, the treatment, the grounds and the dates. California's own category taxonomy is then resolved by lookup against the state's filings, not by the model.
  2. The filing deadline is computed from the statute, with every rule citing its provision: HSC § 1374.30(k) and (j)(3), Civ. Code § 14, § 10, § 7, Gov. Code § 6700. No model touches the date, because a hallucinated deadline is the one error that cannot be recovered from.
  3. It retrieves how California actually decided comparable denials, from the 22,090 Independent Medical Review determinations the state published between 2016 and 2024, each one made by a state-contracted physician reviewer with the reasoning attached.
  4. Claude Haiku 4.5 on Bedrock drafts the appeal, arguing only from what those reviewers found persuasive.
  5. It stops. A Strands interrupt hands the exact letter to a person. Nothing is filed without approval, and nothing is filed by the software after it either: DMHC has no API, so approval produces the finished application with the due date and DMHC's real channels (online, fax, mail), and the person sends it. The page says "that last step is yours" on screen.

How I built it

Strands Agents SDK, used as the spine of the product rather than as a wrapper around one model call. What is in the judged path:

  • One Agent on BedrockModel (Claude Haiku 4.5, temperature 0.1 after measuring behavioural drift at 0.3: at the higher setting the agent sometimes drafted the letter and offered to file "later" instead of calling the tool, which loses the gate entirely).
  • Three @tool functions: compute_filing_deadline and find_precedent are deterministic and the model must argue only from what they return; file_appeal is @tool(context=True) and calls tool_context.interrupt with the exact letter.
  • The gate is a real Strands interrupt: the run returns with result.interrupts populated, a person decides on the page, and the run resumes with an interruptResponse. The same resume works from the terminal (run_agent.py).
  • BeforeToolCallEvent and AfterToolCallEvent hooks drive the page live: every stage and every tool result streams to the letter as it happens, and the callback_handler streams the draft token by token. The grounding eval uses the same AfterToolCallEvent hook to capture exactly what the tools returned, so the letter is checked against the record, not against a guess.
  • The ablation runs the identical Agent with a tool removed from its tools list, so what each tool contributes is measured on the same agent, not on a rewrite.
  • Amazon Nova Lite reads the denial letter through Bedrock before the agent starts; California's taxonomy is then resolved by lookup, not by the model.

Two deterministic engines the models are not allowed near: the statute clock and the taxonomy lookup.

The corpus is public data from California's Department of Managed Health Care: 42,749 published determinations from 2001 to 2026, fetched by script. Splits are temporal and leak-controlled. Precedent comes only from 2016 to 2024; the 3,276 held-out cases are 2025 to 2026, with the reviewer's narrative stripped because it states the verdict.

The same agent is deployed on Amazon Bedrock AgentCore Runtime (ARM64 container built by CodeBuild), so it works in the background rather than on a laptop. Invoked with a real published denial it returns in about 14 seconds with the intake, the statutory deadline and the drafted appeal stopped at the gate. The web surface is also live on AWS App Runner (the demo link on this page), reaching Bedrock through an instance role and capped per day because every run is a real call. It is a single page: the insurer's typeset letter, marked up by hand in a litigator's ink, with a signature line for the gate, because you sign an appeal to file it.

What I measured, and who graded it

Every number below was graded by something I did not write.

Deadline to a drafted appeal waiting at the approval gate: 15.1 seconds median, gate reached 5 of 5 across five different real denials.

Nova reading the denial letter, scored against the category California itself assigned to that case: grounds 30/30, denial date 30/30, treatment category 25/30 exact, diagnosis category 20/30 exact and 26/30 allowing the state's own synonyms.

Grounding, checked mechanically against what the tools returned and with a negative control that must catch seven planted fabrications: 18/18 numeric claims and 12/13 named authorities traced to the record across 20 letters, 7/7 fabrications caught. The one miss is named in the README.

Retrieval coverage across all 3,276 held-out denials: 91.5% get a published overturn rate for their situation, 84.1% get a usable exemplar to argue from, 72.4% at the tightest match. The published rate ranges from 5.0% to 98.6%, which is why the agent reports it rather than predicting an outcome: outcome prediction was measured on day one at +0.005 over a constant and dropped.

Whether those exemplars are about the case, and not only superficially similar, is graded by the reviewer: for each held-out denial, the retrieved exemplars' reasoning is compared with the state physician's own findings for that case, which the agent never sees, against random cases and same-category random cases from the same pool. Retrieved exemplars score 0.245 mean similarity to the reviewer's reasoning against 0.093 for a same-category random draw and 0.035 for random, and beat the same-category draw in 87.8% of 2,754 cases. In the other 12.2% taxonomy matching did no better than category-random, almost all of it at the coarsest tier the floor allows, which is where semantic retrieval would go next.

The ablation

Same model, same six real denials, one tool switched off at a time.

Without California's record, the model invented overturn rates of 40%, 50% and 60% in half the letters, with nothing to cite. With it, 12 of 12 rate claims trace to the state's published figure.

Without the statute engine, on the timeline where the plan sat on the grievance for 120 days, the model counted six months from the plan's answer and landed 86 days late in 5 of 6 cases. That is the appeal forfeited. It was told the rule in words and still did not apply it.

Challenges I ran into

Outcome prediction was the original headline and it was dead: the held-out overturn base rate is 71.89% and precedent lookup scores 0.7234, a lift of +0.005. The product had to be a reporter, not a predictor.

The first live run quoted a pharmacy-wide overturn rate as though it were specific to psoriasis biologics. The tool now returns a plain-English description of exactly what population the rate covers, and the prompt forbids attaching a broad rate to a specific condition.

Asking Nova for California's category directly scored 30%. Reading the misses showed the state files Speech Therapy under Autism Related Tx and Arthritis under Immuno Disorders, classifying by patient context. That is not inference, it is a lookup, so the boundary moved: the model reads, the state's own filings classify. Diagnosis went from 30.0% to 66.7% and treatment from 33.3% to 83.3%.

A Nova embeddings index for the 15.9% of denials with no close taxonomy match was built and abandoned on throttling. The script stays; the README says the index is not built.

Accomplishments that I'm proud of

eval/check_claims.py reads every measurement report and fails the build if the README quotes a number the data does not support. On its first run it found seven disagreements, two of them real. It runs under pytest, and it now audits the gallery captions on this page too.

The ablation above. Every entry at a sponsored hackathon says its tools were essential; I switched them off one at a time and measured what each one contributed.

The page used to stamp FILED after approval when nothing had been submitted, because DMHC has no API. I caught it before recording the demo and fixed the product rather than the wording: approval now hands over the finished application, the due date and the real channels, and the page says "that last step is yours". The tagline says "ready to sign" instead of "filed" for the same reason.

What I learned

A model told the rule in words still gets the date wrong. Given "the clock starts thirty days after the grievance, whether or not the plan answers", Haiku counted from the plan's answer in 5 of 6 late-answer cases. Anything where a wrong answer cannot be recovered from belongs in code that cites its provision, and the model should be handed the result.

The boundary between model and lookup is an empirical question, not a design taste. Nova reads a letter well and classifies it badly, because California's categories encode patient context that a letter does not contain. Moving that one step from the model to the state's own filings doubled the score.

A sentence must never be more real than the artifact behind it. "Filed" was one word, and it would have put a claim on the judged surface that the code did not back.

Reporting beats predicting when the honest lift over a constant is +0.005. The published rate for a matched cohort runs from 5.0% to 98.6%, and telling someone that number, with its exact scope, is worth more than a guess dressed as a forecast.

What's next for Day Thirty

The astronomical holidays in Gov. Code § 6700 (Lunar New Year, Diwali) are supplied as data and the tables are empty; every result says so. Semantic retrieval for the taxonomy gap. Other states' clocks. And the last step: DMHC publishes no API for IMR applications, so today approval hands the person the finished package and the channels, and they send it. If DMHC ever exposes one, that is where the gate's "yes" would go.

Built With

Share this project:

Updates

Submission history