Inspiration
Think about the mail you got this month.
Some of it was a bill- a demand, with a due date, loud if you ignore it. Nobody forgets bills. Bills chase you.
Some of it was the other kind. A letter saying a claim was processed at a lower rate than you expected. A recall notice. A client asking "can we also add the mobile screens?" A price that dropped after you'd already bought.
None of those demanded anything. That's the point. Each quietly opened a window during which you could act, and each closes on its own. When it closes, nothing happens. No notice, no penalty. The right simply stops existing, and the money stays with whoever already has it.
Then I found the number that turned this from an annoyance into a project:
| denied | appealed | overturned | |
|---|---|---|---|
| HealthCare.gov marketplace, 2024 | ~85,000,000 claims | 262,982 — under 1% | 34% |
| Medicare Advantage prior auth, 2024 | — | 11.5% | 80.7% |
(KFF, 2024)
Four out of five appealed prior-authorization denials are reversed and almost nine in ten are never appealed. A denial that would be overturned on request stands permanently because nobody asked.
That is not a legal problem and not a merits problem. It is a noticing problem - which is the only kind an agent that reads your documents can actually solve.
The asymmetry underneath it is what makes this worth building: the counterparty has software tracking these clocks, and you have your memory. An insurer knows to the day when your appeal window shuts. A landlord's management company knows the statutory deadline and knows you missed it. There is no system on your side, so the default outcome is that you lose quietly, repeatedly, in small amounts.
What it does
Lapse reads incoming documents and asks one question of each: what right did this silently start a clock on, and when does it expire against me?
The non-obvious part is that six unrelated-looking problems are structurally one object:
a right you hold + an expiry + a counterparty who benefits from your silence
| What arrives | The clock it starts | Who wins if you forget |
|---|---|---|
| Insurance claim denial | 180-day internal appeal (ERISA) | Insurer keeps the money |
| Client email expanding scope | SOW objection window | Client gets free work |
| Landlord ignores a repair request | Statutory response clock | Landlord |
| Product recall notice | Free-remedy window | Manufacturer |
| Price drop after purchase | Price-protection window | Retailer |
So Lapse isn't an insurance tool or a contracts tool. It's an agent for a primitive, not a vertical.
And it is silent by design. There's no dashboard and no feed. On most days it outputs nothing. In the demo it examines six documents, finds six expiring rights, and surfaces exactly one showing you why it suppressed each of the other five, with a different reason for each.
What surfaces looks like this:
Appeal the denial of the MRI claim closes 2026-09-25, 11 days. Worth $2,340. Keystone will argue the scan was investigational under §7.2 of your Certificate of Coverage. That fails: Medical Policy Bulletin MP-2026-03 reclassified CPT 72148 as medically necessary effective March 1, eighteen days before your scan. Letter drafted. Nothing has been sent.
Note what the demo proves: the most urgent clock (2 days left) and the most valuable one ($1,850) are both suppressed. Lapse is not sorting by urgency or by value it is reasoning about consequence.
How we built it
Five stages per document - detect → date → challenge → rebut → triage with typed Pydantic objects crossing every boundary.
1 · DETECT - a Strands agent with 6 tools reads the document, decides which kind of clock it opened, and goes and reads the governing instrument when the window is set by contract rather than statute. The 180-day appeal period lives in the Certificate of Coverage, not in the denial letter, so the agent has to go find it.
2 · DATE - pure Python, no model. Ask an LLM when a 180-day window opened on 2026-03-29 closes and it answers confidently and wrongly. The model decides which clock applies; Python decides when it closes. No date that reaches a user was produced by token generation.
3 · CHALLENGE - a second agent, in an isolated context with an opposed objective, plays the counterparty and tries to defeat the claim, using the document its own side issued.
4 · REBUT - a third agent answers or concedes, and drafts the letter to send.
5 · TRIAGE - a fourth agent sees the whole surviving docket at once and spends a strict interruption budget, followed by a deterministic Python guard that escalates any valuable right with ≤3 days left regardless of what the model concluded.
Plus a ledger giving it memory between runs (dismiss / snooze / acted, checked before any model call so a muted clock costs nothing), and telemetry so every run shows its own work.
Stack: Strands Agents SDK 1.55.1 · Amazon Bedrock (us.amazon.nova-pro-v1:0, us-west-2) ·
Pydantic · Python 3.13 · 131 offline tests.
Challenges we ran into
The adversary was arguing blind. It received only a summary of the claim never the denial letter its own side had issued, so it couldn't cite the insurer's stated reason. Passing the source document through is what made it argue §7.2, and made the advocate go find the policy bulletin that defeats it. That one fix is the difference between a demo and a toy.
A model argued a claim was "out of time" with 11 days left and the advocate conceded. Timeliness isn't a matter of opinion when Python has already computed the window. A challenge asserting the window has run against a clock computed as LIVE is now void, and the advocate never sees it. But it voids only the timeliness point, because throwing away the whole challenge would silently discard a valid exclusion argument bundled with a bad one.
lapse doctor was a probe that could not fail. It reported a healthy Bedrock provider,
exit 0, and the very next inference call died with ResourceNotFoundException. It was checking
that credentials existed, never that the model would answer. It now makes a real call, and I
validated it against both a working model and a known-blocked one, because a check that can only
pass isn't measuring anything.
Measuring cost took three attempts. The first instrument double-counted (accumulated_usage
is cumulative per agent); the second under-counted 2.7× by keying a baseline on id(agent),
which CPython recycles after GC. The second looked completely plausible. What exposed it was the
agent count, not the token count.
Accomplishments that we're proud of
"Only surfaces when there's a real decision" is structural here, not a prompt instruction. Most agents put don't bother the user unnecessarily in a system prompt and hope. Lapse makes a claim earn its way through: surviving the counterparty's best argument is a test. Being told not to bother someone is not.
The isolation is enforced in code. Separate Agent, separate conversation, opposed
objective, constructed fresh per claim. Two agents reading the same evidence under the same
framing converge, and their agreement carries no information, it's consensus wearing the costume
of corroboration. The run trace shows 19 invocations across 19 agents, a clean 1:1, which is
independent evidence the isolation is real.
No date is ever generated. 131 tests, including regressions pinning every corpus deadline and a mutation test confirming the holiday table actually bites.
It costs $0.028 per document - about $0.84/month for someone receiving 30 documents. Reproduced across two runs agreeing to 3%, because one run gives a value and two tell you whether it means anything.
What we learned
A missed clock is silent, and that is the hardest problem in the product. Six runs of the
identical corpus with detection pinned at temperature=0.0 returned 4, 5, 5, 5, 6 and 6
clocks, a 50% spread. Temperature 0 constrains sampling, not tool-use paths or structured-output
retries. An agent that says nothing because there was nothing to say and one that says nothing
because it failed to look are indistinguishable from the outside, and silence is precisely
what the user is being asked to trust. This is stated plainly in the README rather than buried;
detection recall is the first thing an eval harness should measure.
Guarding most paths is not guarding. Every defect I found took the same shape a check applied everywhere except the newest code. The escalation guard treated an unquantified claim as worth $0. The ledger keyed identity on a model-generated filename, defeating the exact duplicate case it existed for. My own telemetry hook missed the drafter agent.
Where the legal risk actually lives. Nippon Life v. OpenAI pleads unauthorised practice of law against an AI system for autonomous document drafting a closer analogue to this mechanism than the DoNotPay action, which was a deception case. Detecting is safe; drafting is the exposure. For eviction and debt, the defensible line is to surface the deadline and leave the drafting alone.
And a reminder can itself cause the loss. FCBA and Reg E run to arrival at the counterparty, not sending. Telling someone "3 days left" on an arrival-anchored right quietly advises them into a missed deadline. Those need a computed mail-by date.
What's next for Lapse
- An eval set - 40–60 labelled documents, measuring detection recall and date exactness separately. This is the only item that matters until it exists.
- Deadlines as datetimes, not dates. Medicare NOMNC is "by noon of the day following receipt." Storing 1 day is twelve hours too late, at a hospital discharge.
- Clocks that run against the counterparty. Nobody tells you when a landlord blew the 21-day deposit deadline in California that silence is worth twice the deposit. Least served by any existing product.
- Chained clocks. COBRA's premium window anchors on the election date, which hasn't happened when the notice arrives. A clock needs to spawn a successor.
- Real mailbox ingestion, a scheduled quiet run on EventBridge, and AgentCore Runtime.
- Weight the roadmap toward statutory rights. Earny and Paribus automated price-protection so effectively that issuers deleted the benefit. Instrument-defined rights can be revoked by whoever wrote them. Statutory ones cannot.
Built With
- amazon-bedrock
- amazon-nova
- amazon-web-services
- boto3
- litellm
- pydantic
- pytest
- python
- rich
- strands-agents
Log in or sign up for Devpost to join the conversation.