Inspiration
Getting money back is rarely one simple task. It means finding receipts, locating the correct policy, explaining what happened, submitting evidence, monitoring the response and following up—sometimes repeatedly.
Busy consumers often abandon valid refunds and warranty claims because completing the administrative loop takes more time and attention than they can spare.
We built ClaimBack to take ownership of that routine work without taking ownership of the user’s decisions. The agent gathers evidence, prepares the claim, submits within an explicitly authorized scope and follows up. When a merchant proposes a partial settlement or another consequential compromise, ClaimBack stops and asks the person to decide.
What it does
ClaimBack provides consumers with an AI-powered recovery desk for managing refunds and warranty claims.
The working demonstration contains two complete synthetic cases.
Refund recovery
A fictional travel company cancels a trip for which the customer paid $1,408.
ClaimBack:
Collects the receipt, cancellation confirmation and relevant policy. Checks the synthetic claim’s eligibility. Produces a source-backed claim. Submits the request after receiving authorization. Detects that the merchant offered only $880. Calculates and surfaces the $528 shortfall. Stops before accepting the settlement. Allows the user to authorize a policy-backed challenge or explicitly accept the partial amount. Confirms the simulated full refund after the challenge. Warranty recovery
A fictional $349 air purifier fails within its 365-day warranty period.
ClaimBack:
Collects the receipt, fault log and warranty policy. Verifies synthetic eligibility. Assembles and submits the warranty package. Advances the demonstration by a simulated three-business-day period. Performs a routine follow-up. Checks the merchant response and confirms the simulated replacement.
Replacement value is shown separately from cash recovered so the system does not inflate recovery results.
Every case includes readable evidence, a source-backed draft, agent activity records and a downloadable audit package. All merchants, evidence, policies, monetary amounts and outcomes in this demonstration are synthetic. ClaimBack does not contact real merchants or move real money.
How we built it
ClaimBack’s orchestration layer uses the Python Strands Agents SDK with six purpose-built tools:
Gather evidence Assess eligibility Draft the claim Submit the claim Monitor claim status Send routine follow-up
The tools are bound to one case, execute sequentially and enforce authorization and state transitions in application code rather than relying only on model instructions.
The agent deliberately has no tool that can accept a settlement, waive a user’s rights or authorize a compromise. Those actions are handled through a separate human-decision endpoint using current decision IDs, explicit acknowledgement and replay protection.
For a reproducible public demonstration, ClaimBack uses a scripted DemoModel that drives the real Strands event loop without requiring credentials or paid inference. A configurable Amazon Bedrock model can use the same guarded tools for model-driven orchestration, although Bedrock model inference was not used or verified in the public synthetic demonstration.
The application uses:
FastAPI for its API and browser application. SQLite for browser-scoped synthetic state. Random HttpOnly session cookies for session isolation. Optimistic concurrency controls. SHA-256-linked evidence and audit events. A responsive HTML, CSS and JavaScript interface. pytest integration and policy-regression tests. Vercel for the public web demonstration.
A standalone synthetic Strands agent is also deployed separately to Amazon Bedrock AgentCore Runtime in us-east-1. Its verified cloud invocation returned HTTP 200 and produced the expected $528 shortfall decision. The public Vercel interface does not currently call that separate IAM-authenticated runtime.
Challenges we ran into
The hardest problem was not generating a claim—it was defining the agent’s authority.
Routine recovery steps should happen autonomously, but an agent must never convert a merchant’s offer into a customer’s consent. A partial refund could include a waiver, compromise or loss of further rights.
We addressed this through:
Separate human-decision endpoints. Explicit authorization scopes. Current decision identifiers. Replay protection. Required acknowledgement before accepting less than the requested amount. No settlement-acceptance capability inside the agent toolset. Clear separation of simulated cash recovery and non-cash replacement value.
We also wanted the demonstration to be reproducible without hiding its limits. The offline scripted model exercises the real Strands tool loop, while every interface clearly identifies the data and merchant interactions as synthetic.
Accomplishments that we’re proud of Built two end-to-end recovery workflows using the same Strands orchestration and policy layer. Created a genuine human-in-the-loop boundary for consequential financial decisions. Made the agent detect and explain a $528 discrepancy instead of simply reporting a merchant response. Produced evidence-linked claim drafts and downloadable audit records. Separated cash recovery from non-cash replacement value. Deployed the working synthetic web demonstration publicly. Deployed and successfully invoked the standalone agent through Amazon Bedrock AgentCore Runtime. Added tests for both workflows, authorization enforcement, approval replay rejection, duplicate prevention, expired warranties, session isolation, concurrency conflicts and audit mutation detection. What we learned
The most valuable output of an agent is not another answer or summary. It is a completed administrative step that can be verified.
We also learned that a short, evidence-backed decision card is often more useful than a long conversational transcript. When the customer must choose whether to accept a partial settlement, the system should clearly explain:
What was originally requested. What the merchant offered. The value of the shortfall. What evidence supports a challenge. Which rights or opportunities may be lost by accepting.
Autonomy becomes more trustworthy when the boundary between routine execution and human authority is visible in both the interface and the code.
What’s next
The next phase is to pilot ClaimBack with consenting consumers and add:
Authenticated receipt and email ingestion. Durable cloud storage and user accounts. Live merchant and claim-system integrations. Policy extraction with citations and confidence controls. Background claim monitoring and scheduled follow-ups. Encryption, retention and evidence-deletion controls. Idempotency contracts for external actions. Reimbursements, insurance claims and additional recovery categories.
We will measure case-completion rate, active user time, number of interruptions, agent error or overreach rate and actual recovered value before making claims about savings or recovery performance.
Testing instructions for judges Open the live demonstration. Select the Northstar Travel case and click Review & start. Review the authorization scope and select Authorize & start simulation. Inspect the $880 partial offer and $528 shortfall decision. Select Authorize challenge, followed by Run agent. Confirm the simulated $1,408 recovery. Open the Forma Home warranty case. Authorize the simulation and then select Check response. Confirm the simulated replacement. Review Evidence vault and Agent activity to inspect the evidence, tool activity and audit history.
No AWS account, API key or paid service is required to evaluate the public synthetic demonstration.
For local testing:
git clone https://github.com/LightLLM/claimback.git cd claimback python -m venv .venv python -m pip install -r requirements.txt python -m u
Log in or sign up for Devpost to join the conversation.