Inspiration

When a buyer tells PayPal a parcel never arrived, a small seller has about ten days to answer. The proof usually exists: a tracking number, a carrier scan with a signature, an email thread. But it's scattered across PayPal, the store, the carrier and the inbox, and writing it up takes an hour a small shop doesn't have. So sellers either ignore the case and lose money they were owed, or fight cases they can't win and lose a fee on top. We wanted a desk that does the paperwork, tells the truth about which cases are worth fighting, and never invents a fact in front of a PayPal reviewer.

What it does

  • Reads every dispute from the PayPal Disputes API: reason, amount, stage, deadline, and the HATEOAS links that say which responses PayPal will accept right now.
  • Gathers the records behind it as lettered exhibits:
    • from PayPal: the order, the payment with the issuer's address and security-code checks, the tracker, and any refund;
    • from the store: the carrier's scans, the inbox, the listing, the policy shown at checkout, and the customer's history.
  • Argues the case. GPT-5.6 on Azure OpenAI tallies weighted points for each side and recommends fighting or refunding. It writes the response one sentence at a time, and every sentence names the exhibits that prove it.
  • Checks every sentence. Each amount, date, time, tracking number and id must appear in the exhibits that sentence cites. A date written with a time must match one record's date and time together. Anything else is struck before filing, and the brief shows why.
  • Files the seller's choice in one click. Either provide-evidence, with the response, the tracking or refund details and a Bates-stamped PDF of every exhibit, or accept-claim for a full refund. In the sandbox it can then ask PayPal's simulator for a ruling.
  • Shows it all on a desk (AG Grid) of live sandbox cases, and on a landing page where a real case is argued on screen: the letter writes itself as each exhibit lifts out of a tabbed stack.

Judges can try it with no sign-up. Pick one of five stories and Exhibit opens a real sandbox dispute against a demo shop: two cases it should fight, one that's a judgement call, and two it should refund. Exhibit never sees which story you picked, only the records.

How we built it

  • PayPal (sandbox):
    • Orders v2: one-call card checkout with PayPal's test cards, and trackers.
    • Payments v2: captures and refunds.
    • Disputes v1: list, get, provide-evidence (multipart, with a PDF), accept-claim, and the sandbox adjudicate.
    • The sandbox process-chargeback endpoint, so cases are opened by script with no human. We mapped which Visa and Mastercard reason codes become which PayPal reasons: 13.1 and 4855 for not received, C2 for not as described, 4837 for unauthorised, 4834 for duplicate.
    • A cron keeps two cases per story opened ahead, because PayPal reviews a new chargeback for five to eight minutes before the seller can respond.
  • AI: GPT-5.6 on Azure OpenAI through the Responses API, with one forced function call (write_brief) whose schema only allows the moves PayPal's links permit for that case. A rules engine argues from the same exhibits through the same checks when the model is unavailable.
  • The checker: about a hundred lines of TypeScript with twelve unit tests. The win odds come from the tally's weights, not from the model.
  • App: Next.js 15 and React 19, AG Grid Community 36 for the desk, pdf-lib for the Bates-stamped PDF, and a private Vercel Blob store for case records. Deployed on Vercel, with CI on GitHub Actions.
  • Design: the materials of litigation paperwork. The blue backing sheet of court filings, yellow and blue exhibit stickers for buyer and seller, Bates numbers, and a red ink stamp. The hero is a CSS 3D case file on a dark desk.

Challenges we ran into

  • Creating disputes programmatically.
    • The buyer-side create-dispute call needs consent we couldn't script.
    • Card payments failed on a sandbox business account registered outside the US.
    • The fix was a US sandbox business account, one-call card checkouts, and the sandbox chargeback simulator, then mapping reason codes by experiment.
  • Chargebacks only allow two moves. You can fight or refund in full; partial refunds are blocked while a chargeback is open. We changed the model's schema so it can only recommend what PayPal's links allow.
  • provide-evidence takes only one evidence entry unless each names an item. We file the strongest type with everything attached.
  • Keeping the model honest about dates. A sentence can quote a real date and a real time that come from two different carrier scans. The checker now requires date-and-time pairs to match a single record.
  • Showing a five-minute PayPal review without making judges wait. A house pool on a cron opens cases ahead.

Accomplishments that we're proud of

  • Five out of five correct calls on the sample stories, with every sentence surviving the checker.
  • A dispute is opened, argued, checked, filed and ruled on entirely through PayPal's APIs, with no human in the sandbox UI.
  • The pinned brief: you can see which record backs which sentence, and the PDF that reaches PayPal has the same exhibit marks.

What we learned

  • PayPal's dispute lifecycle in detail: inquiry vs claim, which HATEOAS links appear in which state, and how long the sandbox takes to hand a case to the seller and then open it for a ruling.
  • Structured output plus a deterministic checker beats prompting a model to "be accurate".
  • Knowing when to refund is half the value. Sellers lose money on hopeless fights, not just on ignored cases.

What's next for Exhibit

  • Connect a real store (Shopify, WooCommerce) and real carrier tracking APIs in place of the demo shop.
  • Use PayPal webhooks (CUSTOMER.DISPUTE.CREATED) so a brief is waiting before the seller opens the case, plus seller-written house rules ("refund anything under $15 with no tracking") for autopilot.
  • Inquiry-stage tools: messages and partial-refund offers, which PayPal allows before a case becomes a claim.
  • Win-rate tracking per reason and per evidence type, so the odds learn from real outcomes.

Built With

Share this project:

Updates

Submission history