ClaimBack: The Background Agents That Finds Money You're Owed But Never Claimed


Inspiration

People lose real money constantly, not through carelessness, but through time. A repair that was actually covered under warranty gets paid out of pocket because nobody has time to dig up the receipt and read the fine print. A subscription quietly raises its price and nobody notices. An employer reimbursement sits unclaimed until it expires. The problem isn't that people don't care about this money, it's that finding and claiming it is tedious, so it never gets prioritized.

None of this is negligence. The money is recoverable, it's just buried under paperwork nobody wants to audit.

We didn't want to build another chatbot you have to remember to ask. We wanted an agent that works the way a good assistant would: silently, in the background, and only speaks up when it's actually found something and needs one thing from you to finish the job.

What it does

ClaimBack is a background-first financial recovery agent. It watches the documents you give it — receipts, warranties, insurance policies, bills, subscription notices — and continuously looks for money you're entitled to but haven't claimed:

  • "You paid $420 for this repair. Your warranty appears to cover it — here's the claim, ready to send."
  • "This subscription just increased 40%. Your plan allows cancellation within 30 days, here's the annualized cost of staying."
  • "These two charges from the same merchant, 40 seconds apart, look like a duplicate. Here's a dispute draft."
  • "Your insurance policy appears to reimburse this expense — here's the form and the deadline."

Instead of stopping at "you could probably claim this," ClaimBack does the work: it finds the opportunity, verifies it against real evidence (not a guess), checks what supporting documents already exist in your vault, identifies exactly what's missing, drafts the claim or dispute, and asks for a single one-tap approval before anything is ever submitted on your behalf.

There is no chat box. You don't prompt ClaimBack. It runs quietly, and the UI is a calm dashboard, not a conversation: silent when there's nothing to report, and specific the moment there's real money on the table.

How we built it

ClaimBack is deliberately a hybrid system: Strands agents handle reasoning and judgment; deterministic Python code owns every state transition, score, and side effect. No model output is ever trusted to move money, submit a claim, or change an opportunity's status, it can only propose.

ClaimBack architecture flow

Architecture: document ingestion → Strands agents → deterministic workflow → human approval → delivery & follow-up.

Seven narrowly-scoped Strands agents, each with the smallest possible tool inventory for its job:

Agent Responsibility Tools it can touch
DocumentParserAgent Extracts merchant, dates, amounts, line items from OCR text none (pure extraction)
PolicyMatcherAgent Proposes candidate coverage matches for a new expense find_matching_policies
EvidencePlannerAgent Checks the vault for existing evidence, flags what's missing check_evidence_completeness
ActionDrafterAgent Writes the claim email / dispute letter none (read-only drafting)
HumanGateAgent Turns a verified match into a one-tap approval card none (summarization only)
SubscriptionAuditorAgent Reads price-change notices, computes annualized impact detect_price_change_candidates, record_ledger_event
DuplicateChargeAuditorAgent Flags near-duplicate transactions, drafts a dispute detect_duplicate_candidates

None of these agents can send an email, submit a form, or change an opportunity's status. That authority lives entirely in code:

  • A deterministic eligibility service scores every candidate match on identity match (model/serial number), coverage-window validity, and document confidence — a claim only becomes actionable above a fixed confidence threshold, never on the model's say-so.
  • An explicit opportunity state machine — DETECTED → VERIFYING → NEEDS_EVIDENCE → READY_FOR_APPROVAL → APPROVED → SUBMITTING → SUBMITTED → FOLLOWUP_PENDING → RESOLVED — is enforced server-side. The submission workflow independently re-checks that the opportunity is APPROVED, that a human approval record actually exists, and that every evidence requirement is SATISFIED before a claim can be executed. Skip a step, and the workflow raises rather than proceeds.
  • Only after that gate opens does the submission workflow create and execute a SubmissionJob, which a tracking service hands off to a follow-up task for later reconciliation.

Built on AWS end-to-end:

Service Role in ClaimBack
Amazon Bedrock (Claude Sonnet) Reasoning model behind every Strands agent
Strands Agents SDK Agent orchestration with per-agent, least-privilege tool access
Amazon Textract OCR for scanned receipts, warranties, and policy PDFs
Amazon S3 Durable document vault, presigned preview URLs
Amazon DynamoDB (single-table design) Opportunities, documents, coverages, approvals, submissions, follow-ups
AWS Lambda + FastAPI Stateless API handlers for ingestion, decisions, and submission
Amazon EventBridge Scheduled follow-up sweeps so claims don't go quiet after submission

Because ClaimBack's entire job is ingesting documents from the open world (forwarded emails, scans, PDFs), we treated that as an intentional attack surface from day one: extracted document text is passed to agents strictly as data, never blended into instruction context, and no document-reading agent has a submission or external-communication tool in its inventory. Even if a malicious document tried to inject an instruction, there is nothing downstream it could reach.

Challenges we ran into

The biggest design tension was resisting the temptation to let one large agent "figure it all out." Early versions handed too much authority to the model, including eligibility and submission decisions, exactly the pattern that breaks in production the moment a tool call fails silently or a document is malformed. We rebuilt around a strict split: agents propose, deterministic code decides, and nothing external ever happens without a persisted human approval.

Getting the eligibility scoring right was its own problem. A warranty claim can't be "probably fine" — it needs to weigh document confidence, exact model/serial identity matching, and whether the repair date actually falls inside the coverage window, then combine those into one auditable confidence score with a hard cutoff, instead of asking an LLM "does this look covered?" and hoping the answer is consistent twice in a row.

We also had to enforce the approval boundary against ourselves, not just in theory. Early on it was possible to call the submission endpoint on an opportunity that hadn't been approved yet. We closed that by making the workflow independently re-verify status, approval records, and evidence completeness at execution time, not just at the point the button was clicked, so there's no code path, human or automated, that can submit a claim without a real approval on record.

Accomplishments we're proud of

A fully working golden-path demo: ingest a purchase receipt, a warranty document, and a repair invoice, and watch ClaimBack independently surface a verified $420 warranty claim, complete with a drafted claim email, asking for exactly one approval, with no prompting required at any step. Beyond warranties, the same architecture already handles subscription price-hike audits and duplicate-charge detection, proving the "propose → verify → gate → act" pattern generalizes rather than being a one-off trick tuned to a single demo document.

We're also proud of the boundary itself: every agent has a minimal, named tool list, every state transition is enforced in code rather than in a prompt, and the eligibility math is fully inspectable, not a black box a judge has to take on faith.

What we learned

The clearest lesson: users don't want an AI that quietly acts on their behalf, they want an AI chief-of-staff that watches, verifies, drafts, and hands them a one-click decision. Model confidence is not permission to act; a persisted human approval is the only thing that is.

We also learned that agent scope matters more than agent capability. Giving each Strands agent exactly one job and exactly the tools that job requires made the system dramatically easier to reason about, test, and secure than a single agent with a large toolbox ever would have been, and it's the difference between "an AI that drafts your claim" and "an AI that could, in theory, do anything."

What's next for ClaimBack

  • Provider-response ingestion, so a reply from an insurer or merchant automatically closes the loop instead of waiting on a manual follow-up sweep.
  • More evidence templates, covering additional insurers, employer benefit formats, and purchase-protection programs.
  • Amazon Bedrock AgentCore for production-grade runtime isolation as agent execution scales beyond a hackathon deployment.
  • Expansion into more everyday leaks: travel compensation, government benefits, and loyalty/point expirations, all through the same propose-verify-gate-act pipeline.

Built With

strands-agents · amazon-bedrock · claude · amazon-textract · aws-lambda · amazon-s3 · amazon-dynamodb · amazon-eventbridge · fastapi · python · react · typescript

Try it out

AWS Builder Community Posts

  1. Agents for Humans: Architecting an Autonomous Financial Guardian with AWS Strands SDK and Bedrock
  2. Agents for Humans: Using AWS Strands SDK to Build an AI Guardian That Finds Money You're Losing
  3. Agents for Humans: Designing a Tranquil UI with Human-in-the-Loop Security on AWS

Built With

Share this project:

Updates

Submission history