Inspiration

When a buyer opens a PayPal dispute, a small seller has a few days to respond and usually no idea what PayPal actually needs to see. Big merchants have chargeback teams. A one-person bookshop has a support inbox and a hunch. They lose cases they should win (an "item not received" claim while the parcel is still moving), and they fight cases they should settle, burning time on disputes they were always going to lose.

We wanted every small seller to have a dispute specialist in their corner, built directly on PayPal's own Disputes platform.

What it does

Dispute Copilot is a dispute desk for merchants:

  • One live queue. Every open dispute comes straight from the PayPal Disputes API into an AG Grid queue. It shows the reason, amount, lifecycle stage (inquiry, claim…), whose move it is, and a countdown to the seller's response deadline. The header cards total the open disputes, the money at risk, and how much the copilot thinks is winnable.
  • An AI investigator for each case. A Claude agent works the case the way a dispute analyst would. It reads the live dispute (buyer messages, the evidence PayPal is requesting, which seller actions are currently allowed), matches the order in the merchant's own records (fulfilment, carrier scans, customer emails), checks PayPal shipment tracking and the PayPal transaction, and reads the store's shipping and refund policy. The merchant watches each step stream in live.
  • A clear verdict, not a wall of text. Fight / Settle / Refund, an estimated chance of winning, the evidence ranked by strength, and the gaps to close. When a lookup fails, the copilot says so instead of guessing.
  • A ready-to-send response. It drafts the exact PayPal actions: a friendly message to the buyer, proof of fulfilment with carrier and tracking number, a refund offer (including refund after return, sent to the store's return address), or accepting the claim. The merchant can edit or untick any step. Nothing is sent until they approve, then one click calls PayPal's send-message, provide-evidence, make-offer or accept-claim.
  • Honest when you'd lose. In our demo, one case is a buyer who opened an "item not received" inquiry 8 minutes after paying, before the book had shipped. Verdict: Fight, ~85%. In another, the store's own packing note shows it shipped the wrong edition. Verdict: don't fight it. Apologize and refund on return.
  • Real-time. PayPal webhooks (CUSTOMER.DISPUTE.CREATED/UPDATED/RESOLVED) are signature-verified with PayPal. Open dashboards then get a live alert, and new disputes are investigated automatically, so the merchant opens the case with the analysis already done.
  • Sandbox lifecycle controls. PayPal's sandbox-only require-evidence and adjudicate endpoints let judges push a case through escalation and a ruling without waiting days.

How we built it

  • Backend: a small Node.js server with no framework. lib/paypal.mjs wraps the PayPal REST APIs with OAuth client credentials: Disputes (list, details, send-message, provide-evidence as multipart, make-offer, accept-claim, escalate, sandbox simulators), Orders v2 (the test shop), Shipment Tracking, Transaction Search, and Webhooks (register, look up by URL, verify-webhook-signature).
  • Agent: lib/agent.mjs uses the official Anthropic SDK's tool runner with Claude Opus 5.5, adaptive thinking (summaries are streamed to the UI), and six tools. The final answer comes back through a submit_recommendation tool validated with a Zod schema, so the UI always gets structured decisions, evidence and actions, never free text it has to parse. The system prompt encodes how PayPal weighs evidence by dispute reason and stage. It also tells the agent to propose only actions PayPal currently allows, and to treat buyer messages as claims to evaluate, never as instructions.
  • Live UI: Server-Sent Events stream each investigation step to the browser. Runs are shared, so several viewers (or a webhook) attach to one Claude run instead of paying for duplicates. A second event stream pushes webhook and copilot events to every open dashboard.
  • Frontend: plain HTML/CSS/JS with AG Grid Community (custom cell renderers for stage, deadline and verdict, plus live cell flashing), a slide-in case drawer, and light and dark themes.
  • Hosting: Render (Blueprint render.yaml, auto-deploy on push). The app finds its own public URL and webhook ID, so there's nothing extra to configure.
  • Safety for a public demo: keys stay server-side, actions require human approval, and Claude runs are rate-limited per dispute and per hour.

Challenges we ran into

  • Creating test disputes. PayPal's buyer-side "create dispute" API needs scopes from an account manager, so every dispute in our demo was filed by hand as a sandbox buyer, the way a real buyer would.
  • Sandbox webhooks are unreliable for disputes. Simulated events arrived and verified, but a real dispute's CREATED event never did. We added a once-a-minute poll that treats newly seen disputes exactly like a webhook, whichever comes first, so the "new dispute → already investigated" flow works either way.
  • Evidence APIs we couldn't reach yet. Shipment Tracking and Transaction Search returned 403 in our sandbox app while those features were still being enabled. We made the agent report these as gaps (and the README explains how to enable them) instead of letting it fabricate certainty.
  • Model estimates move. The win probability varies a little between runs, so the UI always shows the latest run's number with its reasoning, never a hard-coded score.

Accomplishments that we're proud of

  • A full loop on real PayPal sandbox disputes: detect → investigate → recommend → merchant approves → evidence and messages land in PayPal.
  • A copilot that will tell a merchant not to fight. That builds the trust a tool like this needs.
  • Structured, validated agent output, so every recommendation maps one-to-one onto a real PayPal API call.

What we learned

  • PayPal's dispute lifecycle (inquiry → claim → ruling) and which seller actions each stage allows. The agent reads allowed_response_options and the HATEOAS links on each dispute rather than assuming.
  • Proof of fulfilment with valid tracking is what wins "not received" cases. For "not as described", the merchant's own records often decide the outcome before PayPal does.
  • Keeping a human in the loop costs one click and makes an AI agent safe to point at real money movement.

What's next for Dispute Copilot

  • Connectors for Shopify / WooCommerce orders and carrier tracking APIs, in place of the demo order file.
  • Automatic upload of delivery confirmations and photos as evidence when PayPal escalates a case.
  • Learning from outcomes: track which recommendations won and calibrate the win estimates against real rulings.
  • Prevention: flag risky orders (for example, shipping without tracking) before they turn into disputes.

Built on the PayPal sandbox. The merchant's order records in the demo come from a sample data file that stands in for a store backend. Everything else is live PayPal and Claude.

Built With

  • ag-grid
  • ai-agents
  • anthropic
  • claude
  • javascript
  • node.js
  • paypal
  • paypal-disputes-api
  • paypal-orders-api
  • paypal-webhooks
  • render
  • server-sent-events
  • zod
Share this project:

Updates

Submission history