Inspiration

A chargeback turns a small merchant into an investigator. The answer may be buried in an order, a delivery event, a customer's address-change message, or the merchant's relationship with that buyer. The deadline keeps moving while the owner is running the business.

We wanted to give that owner a prepared decision: the relevant facts, a reasoned recommendation, and a way to act from their phone. Sometimes the right answer is to concede. Rebuttal is designed to make that choice deliberate, rather than treating every dispute as something to fight.

What it does

Rebuttal assembles a case file from order details, shipping records, customer communications, and merchant history and policy. Its agents recommend fighting, conceding, or refunding an inquiry, and prepare the supporting evidence narrative. The console lets the merchant inspect the underlying records and the recorded recommendation.

Before a consequential Stripe action, a separate approval hook checks the merchant's policy. Under the default policy, amounts of at least $200, estimates in the 0.35–0.65 uncertainty band, and any non-fight action request an owner decision. Telegram presents Fight, Concede, or Hold, so the owner can respond away from the desk.

Our 2:58 demonstration follows a $340 Stripe test dispute. It shows the actual console, the order and delivery records, the customer's messages, the recommendation, the real Telegram alert, and the confirmation after the owner chooses Concede on their phone. A fresh Stripe readback and the recorded owner audit confirm that the same test dispute closed as lost, which is Stripe's status for the concession shown. The console displays CONCEDED · CLOSED.

How we built it

AWS Strands Agents SDK runs a seven-node evidence graph and a separate execution agent. Intake reads the Stripe dispute and payment context. Orders and shipping can begin together; communications and history wait for the order lookup to resolve the merchant's customer ID. Strategy waits for all four evidence investigators. The drafter receives both the strategy and the underlying evidence directly, so it does not have to reconstruct facts from a recommendation alone.

Amazon Bedrock provides the model inference; the recorded generation used Claude Sonnet 4.5. Structured outputs, source checks, and the execution hook separate evidence interpretation from action. A recommendation is not permission to execute.

Codex and Google Antigravity assisted with implementation, review, and debugging; Codex also helped produce the edited demonstration. The runtime AI is the Strands/Bedrock pipeline described above. Narration is AI-generated.

The user interface is Next.js, React, and TypeScript. Supabase holds the case and audit records. The phone reply is handled by the existing AWS Lambda and Amazon Bedrock AgentCore integration, which applies the owner's decision through the Stripe test API. The repository also contains Gateway, Memory, deadline-sweep, and deployment components.

The recording combines a local console and local Strands pipeline with real Bedrock inference, a genuine Telegram exchange, and the existing cloud callback. The test case was recovered into the case store before generation. This recording does not establish automatic cloud ingestion or deployment of the latest local console and graph repairs.

Challenges we ran into

The hardest problem was keeping every displayed claim tied to evidence. Early outputs confused planned actions with completed actions, inferred intent from missing messages, and sometimes relied on incomplete evaluator inputs. We tightened source handling, normalized money and timestamps, supplied complete policy context to evaluation, and added checks for unsupported assertions and evaluator false positives.

The graph also exposed a real dependency: communications and customer history cannot run correctly until the order lookup has identified the merchant's customer. We made that ordering explicit and guarded the fan-in before strategy and drafting.

Finally, a successful chat message is not enough to prove an action happened. We checked the resulting Stripe dispute, matching case status, and owner audit separately before using the completed outcome in the video.

A final consistency pass made owner replies update the selected decision across local and cloud stores. Closure memory now requires consistent execution audits; missing or conflicting evidence produces an explicit skip. Local regression tests cover cloud-only decisions, preserved history, failed writes, and schema upgrades.

Accomplishments we're proud of

  • The demo shows inspectable product evidence and an actual owner decision from their phone.
  • The $340 test case closed through the existing phone callback, with matching Stripe and case-store readbacks.
  • The updated submission source passed 256 Python tests locally with synthetic fixtures and offline provider doubles. The unchanged console previously passed 12 focused tests, type checking, and a production build; those console checks were not rerun for this revision.
  • The public video shows the full decision journey in under three minutes, with English captions and clearly disclosed synthetic records and Stripe test mode.

What we learned

The agent's useful output is a decision someone can trust and act on. That requires more than plausible prose: correct data dependencies, visible source records, a separate execution boundary, and a check of the resulting state. We also learned to evaluate the evaluator; a false accusation of hallucination is still a testing defect.

What's next

Apply the prepared database migration, deploy the tested consistency fixes and updated console, and verify automatic ingestion and provider readbacks with the current source. Then add direct merchant-system adapters and test with consenting merchants to measure review time and decision quality.

This is a prototype using synthetic merchant evidence and Stripe test mode. Direct carrier and Gmail adapters are not implemented. Win estimates are not calibrated against real merchant outcomes, and real recovery rates or time savings have not been measured. The recorded case's decision row and closure-memory label remain historically inconsistent; terminal dispute state and the owner audit establish the demonstrated concession. The updated source contains locally tested fixes, but no new deployment or historical repair was performed. The PostgreSQL migration is prepared and has not been applied. Live synchronization remains unverified. The deadline sweep has a separate silence policy, so the approval flow is not an indefinite-wait guarantee.

Source and architecture

Evaluate the pinned submission revision: Source and local setup · Consistency verification · Architecture diagram.

Built With

Share this project:

Updates

Submission history