Inspiration

Agents that can check out for you are arriving. The part nobody wants to trust is the payment credential: a confused or manipulated model with a way to pay can overspend, buy twice, or keep retrying after something broke. Prompts like "never spend more than $50" are advice to the model, not a control. I wanted the control to live in code the model cannot talk to.

What it does

SpendGuard sits between an AI shopping agent and PayPal. Before any money moves, you set a policy:

  • Hard limits: per order, per day, per merchant per day, allowed merchants, currency. Over a limit means denied, and PayPal is never called.
  • A circuit breaker: after N failures in a row the agent is halted until a human resets it.
  • Soft rules that ask a human: orders above an auto-approve amount, and prices that rose more than X% between the quote and checkout.

Approvals clear only the soft rules. The approved proposal is re-checked against the hard limits at click time. The agent has two tools (search_shop, purchase); there is no tool to change the policy, approve, or reset, so a prompt injection has nothing to call. Every decision is written to an audit trail with the reason in plain words.

The web demo has seven one-click scenarios: a normal purchase, an over-limit item, an "ignore your limits" prompt, a purchase that needs approval, a price jump at checkout, a daily limit, and the circuit breaker.

How PayPal and AI are used

  • PayPal Orders API v2 on the sandbox: OAuth2 client credentials, POST /v2/checkout/orders (intent CAPTURE, idempotent PayPal-Request-Id), buyer approval on PayPal's page, then POST /v2/checkout/orders/{id}/capture. Amounts are integer cents. Only the sandbox host is supported.
  • Verified live on 2026-10-02 (log: sandbox_run.md in the repository): token, order creation and capture are real calls. With PAYPAL_BUYER=card, PayPal's public sandbox test card pays and the order is captured while it is created (COMPLETED). The default flow was checked up to the real PAYER_ACTION_REQUIRED response and PayPal's ORDER_NOT_APPROVED refusal of an early capture; the buyer-approval page itself was not driven (no sandbox buyer login was available). Denied and halted proposals produced no PayPal request; the circuit breaker was driven by real INSTRUMENT_DECLINED answers.
  • AI: an LLM drives the shop through tool use (Anthropic Messages API). An offline rule-based planner is the default so anyone can run the demo without a key; the guard behaves identically for both.

How I built it

Python standard library only (no install step). A guard module with atomic reservations (20 parallel purchase attempts cannot overspend), a thin PayPal layer with a sandbox client and an offline replay client behind one interface, a planner interface, and a small local web UI. 62 unit tests (plus 3 that run only against the live sandbox), including deliberate limit breaches; as a check on the tests, I broke four guard rules one at a time and the suite failed each time.

Challenges

  • Keeping the model out of the decision: the agent's tool list has no path to any human control, and the tests include a model that tries to call an invented set_policy tool.
  • Hard versus soft: an approval must not become a way around a limit, so approvals are re-checked against the hard rules when clicked.
  • PayPal's buyer-approval step is a real gate in the Orders flow, so the app shows it (the order waits in PAYER_ACTION_REQUIRED) instead of hiding it.
  • Real sandbox behaviour differed from my hand-written test responses (no purchase_units in the create response, different link hosts, a reproducible PAYER_CANNOT_PAY for exactly 34.99 USD); the replay fixtures were corrected to match.

What I learned

Limits are easy to state and easy to get subtly wrong: boundary cases (exactly at the limit), quantity times unit price, order-splitting, day rollover, and races between parallel calls. Each of these became a test that fails if the rule is removed.

What's next

Per-category limits, a webhook listener to reconcile captured payments against the ledger, support for the PayPal Agent Toolkit / MCP server as the payment tool, and a policy file that can be signed by the user.

Testing instructions

No account, key or install is needed (Python 3.10+, standard library only).

  1. git clone https://github.com/selenium132/spendguard.git && cd spendguard
  2. python3 -m spendguard and open http://127.0.0.1:8765. With no credentials it starts in replay mode (PayPal answers come from spendguard/fixtures/); the header badge says "PayPal: replay".
  3. Click scenarios 1 to 7. In replay mode use "Buyer approves (replay)" then "Capture" for an order waiting for the buyer.
  4. python3 -m unittest discover -s tests -t . runs 62 tests (3 live-sandbox tests are skipped without credentials).
  5. Optional, real sandbox: create a sandbox REST app at developer.paypal.com, copy .env.example to .env, fill in PAYPAL_CLIENT_ID and PAYPAL_CLIENT_SECRET, set -a; . ./.env; set +a, optionally PAYPAL_BUYER=card, then run again; the badge says "PayPal: sandbox". My own run against the sandbox is recorded in sandbox_run.md. No keys or test accounts are stored in the repository.

Tools used and how

  • PayPal developer platform, Orders API v2 (sandbox): creates and captures orders.
  • Anthropic Claude via the Messages API (optional): the agent's brain, through tool use.
  • Claude Code (AI coding agent): used to write the code, tests and documents under the owner's direction.
  • No sponsor tools (AG Grid, APIMatic, etc.) are used.

New project

New. Created on 2026-10-02 from an empty folder, inside the submission period. No pre-existing code.

Built With

Share this project:

Updates

Submission history