Inspiration

We didn't start with prior authorization. We started by looking for the kind of problem that everyone in an industry complains about and nobody has fixed, because fixing it isn't glamorous. Prior authorization kept coming up.

It's the approval an insurer demands before a patient can get treatment that is, in every other respect, already covered. We expected it to be an annoyance. The research says it's closer to a second job. The American Medical Association's surveys put it at 39 to 43 prior authorizations per physician per week, taking up around 13 hours of physician and staff time. That's most of a working day, every week, spent re-typing clinical facts into someone else's form. Roughly 31% of physicians say their requests are often or always denied, and denials have been climbing for five years straight.

Then we found the detail that actually made us build this. Appeals work. Most practices never file them. They assume they'll lose, and every appeal has to be written from scratch by whoever has a spare hour, so it never happens. A denial that should have been a conversation quietly turns into treatment the patient doesn't get.

The tools that do exist are sold to large hospital systems with deep records integrations. A three-provider behavioral health clinic doesn't get any of that. They get a portal login and a fax number, and the hard part is still theirs: does this even need an authorization? does the note actually satisfy this insurer's criteria? how do we argue back?

That's the part we wanted to take off their desk.

What it does

You upload the clinical note the practice has already written. Nothing gets re-typed.

Attest reads it and figures out from the document itself which insurer's policy applies. There's no dropdown, no setting. It then checks the note against that insurer's own published criteria, one criterion at a time, and tells you where each one stands: met, not met, or not enough information to say. Every verdict comes with the exact sentence from the note behind it, highlighted in place in the note. Where something's missing it asks a specific question back to the practice, like which medication, at what dose, for how long, instead of filling the blank in itself.

Then it builds the submission packet and stops. A clinician has to read and approve the clinical claims before anything goes out.

When the denial comes back, which arrives days later as its own separate document, it reads that too. It works out which criteria the insurer is actually contesting, and drafts the appeal using the insurer's own policy wording against them, backed by the quotes from the chart. It tracks the filing deadline. And once a clinician approves an appeal, that approved language comes back as a starting point the next time the same criterion gets contested, so nobody writes the same argument twice.

Then it stops again and waits for a human.

That's the shape of the whole product, really: it does the assembling and the arguing, a person does the deciding and the sending.

How we built it

It's a multi-agent system built on the Strands Agents SDK. There's an intake specialist that reads the note, a criteria specialist that grades it against the policy, a packet specialist that assembles the submission, and an appeal specialist for when a denial shows up, all coordinated by an orchestrator. The orchestrator decides what happens next. It never decides what's true. Anything that could be checked is checked by ordinary code instead.

A few things we committed to early and never moved off:

Insurer policies are data, not code. Each one is a file listing that policy's criteria, quoted word for word from the real published document, with a link back to the source. We shipped two insurers covering the same treatment deliberately, because they want genuinely different things. One wants four failed medication trials plus a failed course of therapy. The other wants two at an adequate dose plus an augmentation trial. Same patient, same treatment, different answer. Later we added an entirely different specialty, physical therapy, to see how much code it would take. It took none. It was one more file.

We built the answer key before we built the grader. We put together a set of realistic synthetic patient notes and worked through them by hand first, writing out criterion by criterion what a careful human reviewer should conclude for each one, and why. Only then did we write the code that does the matching. That gave us something to grade the agent against that we couldn't quietly adjust later. "It works" stopped being a judgment call and became a score.

The evidence check is plain code. Before any quote can appear in a document, it's checked character by character against the source note. Not asked for politely in a prompt. Checked.

The stack under all of that is deliberately small. Strands for the agents, one model provider named in exactly one file, and ordinary Python for the verifier, the gates and the policy routing, which are the three things every safety claim in this product rests on. Strands is the only AWS SDK in the project. There is no Bedrock and no AgentCore, and you don't need an AWS account to build, test or run any of it. That wasn't the original plan. The next section is why.

The reviewer's screen is a web app, deployed publicly so anyone can walk through a case without installing anything. Every note and denial letter in the demo is synthetic. There's no real patient information anywhere in this build, and real records integration is deliberately out of scope for now.

Challenges we ran into

The cloud we planned on never opened, so we stopped planning on it. The intent from day one was Amazon Bedrock with AgentCore. Our AWS account went into verification and every model came back NOT_AUTHORIZED, account-wide. Our credentials were fine, our permissions were fine, the region was fine, and the control plane cheerfully listed 116 models we weren't allowed to call. It wasn't our gate to open. We re-checked it for four days and it never moved.

So we moved to Gemini and put the entire model choice behind a single file, so nothing else in the codebase knows or cares which provider it's talking to. Then, with a day to go, we made the harder call: we took Bedrock and AgentCore out of the architecture altogether rather than leave them in as an intention. The AgentCore entry point we'd already built went with them, and that was real working code with fifteen passing tests. It was deleted not because it was wrong, but because the runtime it was an entry point for was never going to exist in this build. Nothing the product promises went with it, either. The rule that module existed to demonstrate, that no document leaves without a clinician approving a named version of it, was never actually located in that file. It lives in the gate code, and a terminal walkthrough still shows it with no browser and no key.

What forced the deletion, rather than just a quiet edit to the plan, was the dependency list. Our pyproject.toml declared the AgentCore SDK as a hard requirement, and one of our own tests asserted that it did. So a README saying "there is no AgentCore here" would have shipped alongside an install that pulls one into every judge's clone. This entire product exists to reject claims that don't match their source, and we weren't going to suspend that principle for our own packaging. The dependency came out, and the test that used to assert it was present now fails the build if it ever comes back. The architecture claim is enforced by a check rather than by a sentence, which is the same move we make everywhere else in this codebase.

Then the free tier turned out to be 20 requests per day. Per day, not per minute. One full pass over our test cases is about thirty calls. We literally could not run our own tests twice in an afternoon.

Our own bug fix taught the model to lie. The note-reading step kept occasionally coming back with an empty list of procedure codes, so we did the obvious thing and made the field required. Next run, it confidently produced two procedure codes for a note that contained none, filled in from general knowledge of how that treatment is usually billed. We had built a form where "there's nothing here" was impossible to say, so it made something up instead. That's precisely the failure this entire product exists to prevent, and we had reproduced it ourselves, with a fix meant to improve reliability. We took the constraint back out and made the absence something it could report.

Getting "we don't know" to survive all the way to the screen. There's a real difference between the note showing a requirement wasn't met and the note simply not saying. The first is an argument to have with the insurer. The second is a question to ask the practice. If you collapse them, you send people hunting for a document that cannot exist, or you argue a point you're going to lose. Keeping those two apart at every layer, all the way to the wording on screen, took more care than any individual piece of the pipeline.

Saying the right thing the wrong way is still wrong. Some criteria are things that must be absent, like a seizure history for this particular treatment. Our screen was showing a green tick next to "Seizure disorder or any history of seizure," which is technically the correct verdict and reads, to anyone who isn't a clinician, as this patient has a seizure disorder. It means the opposite. A correct answer displayed badly is just a wrong answer with extra steps, and we ended up rewriting how every one of those verdicts is worded.

Accomplishments that we're proud of

It can't make things up, and that's enforced rather than requested. Every quote in every generated document is verified word for word against the source note before it's allowed anywhere near paper. We made it strict enough that it rejects a correct paraphrase, which felt wrong for about a day. Then we remembered that paraphrase is exactly the thing we're defending against. In a document that goes to an insurer about a real patient, "close enough" is how a sentence nobody ever said ends up in someone's file. Across every demo case, every evidence quote in every document traces back to the note exactly.

Twice our answer key was wrong and the agent was right. One of our notes requested a treatment and never actually stated the procedure codes. Two others never stated the patient's age, which is itself one of the criteria. We'd written those notes ourselves, on purpose, trying to make them complete. We fixed the notes rather than lowering the bar. The agent caught documentation gaps its own authors had missed, which is the product's entire pitch demonstrated against the people who built it.

Rationing turned into the best feature we didn't plan. To survive that daily quota we started recording every model response and committing the recordings, so the tests replay them instead of calling out. What that actually produced is a project anyone can clone and run with no API key, no network, no AWS account and no cost: all 380 tests, offline, in about thirty seconds, with the same results we get, on Linux, macOS and Windows alike. The public demo works the same way, which is why it doesn't ask anyone for credentials.

The moment we built the whole thing to show. Upload a note for one insurer, then a note for a different one, into the same untouched screen, and it reaches two completely different sets of criteria, because it worked out which policy applied from reading the document. And when it hits an insurer we have no policy on file for, it stops and says so rather than improvising, because telling a practice "no authorization needed" when the truth is "we couldn't find the rules" is the single most expensive mistake this tool could make.

What we learned

A checker beats an instruction. Asking a model to quote exactly is a hope. Verifying the quote is a guarantee. Every time we found ourselves writing a more emphatic instruction, the better move was to stop asking and start checking. The last thing we pointed that at wasn't the model at all. It was our own dependency list, where a test now fails the build if the SDK we said we removed ever reappears.

Never build a form where "nothing here" can't be said. Not for a model, not for a person. If the only way to proceed is to fill something in, something will get filled in.

"I don't know" is a feature and it needs protecting. It's the answer that's easiest to lose, because it looks like a gap in the product rather than an honest result. It's usually the most useful thing the system can say.

Hard limits made a better product than a bigger budget would have. We only built the offline, reproducible version of this because we couldn't afford to call the model. It's now the thing that makes the project verifiable by anyone.

The interface is part of the safety argument. We spent a long time making sure the system couldn't produce an unsupported claim, and then nearly undid it by displaying a correct result in a way that read backwards. Being right isn't enough if the screen miscommunicates.

A blocker stops being a status and becomes an answer. We spent four days treating an access gate we didn't control as a thing that was about to clear. Deciding that it wasn't cost us an afternoon of deletions and bought back a stack we can describe honestly in one paragraph.

What's next for Attest

Real connections. Right now Attest produces a submission-ready document and a human sends it. The obvious next step is submitting directly through insurer portals and pulling notes from the practice's existing records system, which is mostly a compliance and agreements problem rather than a technical one, and which needs proper handling of real patient data before it goes anywhere near a real clinic.

Track what actually happened. Today, when the system offers previously approved appeal language, it means a clinician signed off on it. It does not mean the insurer overturned the denial, because nothing records the outcome yet. Once outcomes are tracked, the system can prefer the arguments that actually won, which is when the appeal library starts compounding in value.

More insurers, more specialties. Each one is a policy file, not a rewrite. The pipeline that handles behavioral health already handles physical therapy without a line of code changing, and we'd like to see how far that holds.

Close the loop on the questions it asks. When Attest tells a practice something's missing, they should be able to answer it right there and have that criterion re-checked immediately, instead of starting the case over.

Measure the real number. We can show that it produces a complete, evidence-backed packet in seconds. What we can't yet show is how many minutes that saves a real office manager on a real case. That takes putting it in front of actual practices and timing it honestly, and it's the number that would matter most to anyone deciding whether to use this.

Built With

  • bedrock-agentcore
  • css
  • fpdf2
  • gemini
  • github-actions
  • google-ai-studio
  • html
  • markdown
  • pydantic
  • pytest
  • python
  • pyyaml
  • strands-agents
  • streamlit
  • streamlit-community-cloud
Share this project:

Updates

Submission history