Inspiration
Agentic commerce is arriving faster than the safety layer to support it. Large language models can now browse products,compare specifications, and recommend purchases.But the moment you give an AI agent the ability to execute a payment, you introduce a class of risk that did not exist before: -The agent can be prompt-injected through a product description or a review. -The agent can hallucinate a price, a specification, or a discount. -The agent can reinterpret an explicit limit like "under $800" as "approximately $850". -The agent can be manipulated by a third party into overspending on the user's behalf. Existing AI shopping assistants either do not move money at all,or they hand the model unrestricted access to the user's wallet.Neither is acceptable. We built PayGuard AI around a single principle: AI gets intelligence. The user keeps financial authority. The model is allowed to understand. It is never allowed to authorize.That boundary is enforced in code, not in a prompt.
What it does
PayGuard AI is a transaction firewall for AI-powered commerce.It converts a user's natural-language spending rules into a structured, enforceable policy, evaluates any proposed purchase against that policy deterministically, and only then creates a PayPal sandbox payment for human approval.
A complete run:
1.The user types:"I need a programming laptop under $800 with at least 16GB RAM. No refurbished."
2.PayGuard sends that to Groq's openai/gpt-oss-120b model and receives a structured policy:
{
"max_budget": 800,
"currency": "USD",
"requirements": [
{ "text": "at least 16GB RAM", "keywords": ["16gb", "ram"] }
],
"restrictions": [
{ "text": "no refurbished", "type": "forbid", "keywords": ["refurbished"] }
],
"requires_human_confirmation": true,
"risk_level": "LOW"
}
3. The user proposes a transaction — item, price, specs.
4. The deterministic policy engine evaluates it in pure code — no LLM involved — and returns one of:
- ALLOW with a list of satisfied checks, or
-BLOCK with the exact reason: `Price $849.00 exceeds authorized maximum of $800.00 by $49.00.`
5. Only if the decision is ALLOW does Step 3 unlock, offering a Pay with PayPal button.
6. The user approves in the PayPal sandbox.
7. PayGuard captures the order and returns a green receipt with the Capture ID.
8. Every policy, decision, order, and capture is timestamped in an audit trail at `/audit`.
## How we built it
PayGuard AI is built on three deliberately separated layers:
| Layer | Responsibility | Technology |
|---|---|---|
| Reasoning| Understand natural language, produce a structured policy | Groq `openai/gpt-oss-120b` with JSON mode |
| Enforcement| Evaluate the purchase against the policy | `src/policy.js` — pure Node.js, no AI |
| Execution | Move the money, after human approval | PayPal Orders API v2 (sandbox) |
The model has no access to PayPal credentials. The policy engine has no access to the model. The two layers cannot influence each other.
Stack:
-Node.js 24 + Express — the server, running from a single `server.js`
-Groq SDK with `openai/gpt-oss-120b` — structured JSON intent parsing
-PayPal REST API v2 — OAuth 2.0 token, `POST /v2/checkout/orders`, `POST /v2/checkout/orders/{id}/capture`
-Vanilla HTML/CSS/JavaScript— no framework, mobile-first UI
-Termux on Android— the entire project was built and is demoed from a phone
-Render— hosted demo at `https://payguard-ai-qloi.onrender.com`
-GitHub — public repo at `https://github.com/Sule-Bashir/payguard-ai`
Key design decision — typed restrictions.In an early version, the model returned restrictions as plain strings like "Must be new,not refurbished".The policy engine then extracted keywords, found the word `new` in the item text, and produced a false BLOCK— treating a satisfied requirement as a violated restriction.
The fix was to make the model return structured restriction objects with an explicit `type`:
- `{ "type": "forbid", "keywords": ["refurbished"] }` → finding the keyword is a violation
- `{ "type": "require", "keywords": ["new"] }` → not finding the keyword is a violation
The AI layer now expresses intent unambiguously, and the code layer enforces it without guessing.
## Challenges we ran into
1. The Gemini credit card problem. We started with Gemini for the LLM, but Google required card verification to activate the API key. We migrated to Groq's `openai/gpt-oss-120b`, which offers a generous free tier with no card required — and which turned out to be a better fit because it natively supports JSON mode.
2. PayPal sandbox account creation failed for our region.PayPal's developer dashboard would not create a sandbox Business account from Nigeria — the "Create App" button failed with a generic"Something went wrong"error. We worked around this by creating the app through the PayPal sandbox site directly with a US-region profile, which PayPal's own community documentation confirms is the supported path for affected countries.
3. A phantom environment variable.After we set up `.env` correctly, the server kept returning `401 Invalid API Key`. The cause was not the code and not the key — a leftover `export GROQ_API_KEY="gsk_your_new_key_here"` line in `~/.bashrc` was defining the variable as a placeholder string and dotenv refused to override an existing environment variable.The `.env` file was correct the whole time. Removing the line from `.bashrc` fixed it instantly. Lesson:the environment is part of the program.
4. Mobile browser caching hid our HTML updates.For several hours, a button appeared not to work because Chrome on Android was serving a cached copy of `index.html`. We solved it by always loading the page with a version query string (`?v=N`) that changes on every update — a technique that has now become standard practice for the project.
5. Tab reloads lost state during the PayPal round-trip.When Chrome switched to the PayPal tab for approval,it evicted the PayGuard tab from memory.When the user returned, the JavaScript variable holding the Order ID was gone.We fixed this with `localStorage` persistence, so the policy, decision,and pending Order ID survive a reload.
6.`fetch` timeouts on the sandbox API. Node's built-in `fetch` occasionally timed out to `api-m.sandbox.paypal.com` on mobile networks,even though the capture had already succeeded on PayPal's side. We added an automatic one-time retry with a two-second backoff, and hardened the UI to tell the user when a network failure is transient.
## Accomplishments that we're proud of
-A complete end-to-end agentic-commerce flow, proven on real infrastructure.Not a mockup — a $749.00 PayPal sandbox payment was actually created, approved, and captured. The PayPal confirmation email arrived at the sandbox Business account.
-A clean architectural separation between reasoning and authority.The AI can never modify the policy it generated, never sees PayPal credentials, and never decides whether money moves.Every financial decision is made by deterministic code.
-The project was designed, built, and demoed entirely from an Android phone using Termux.No laptop,no desktop environment,no IDE.Every line of code was written in `nano` on a mobile terminal.
-A timestamped audit trail that records every policy, every decision (with reasons),every order,and every capture. The audit page is what turns PayGuard from a demo into something closer to a compliance tool.
-Two prize-relevant tracks in one project.PayGuard naturally targets both Best Use of PayPal + AI and Best Use of Agentic Commerce, plus the Best Use of Render sponsor prize through the hosted deployment.
## What we learned
- Separating an LLM from a financial decision is not just a safety measure — it is a design pattern. Once the model's output is structured, the code that enforces it becomes simple, auditable, and testable.
- Structured output needs to be typed, not just shaped. Returning a restriction as a string forces the code layer to guess intent. Returning it as `{ type, keywords }` removes the guesswork entirely.
- Deterministic code catches LLM ambiguity that no prompt can fix.The false-BLOCK bug was not solvable by writing a better prompt. It was solvable only by changing the *data contract* between the model and the code.
- On mobile, caching is the enemy of iteration.A single stale `index.html` can make a working server look broken for hours. Version query strings solve it.
-The environment is part of the program. A stray `export` in a shell profile can silently override a correctly configured `.env` and produce errors that look like code bugs.
- PayPal's developer platform is remarkably fast when it works.The capture flow, once the sandbox is set up correctly, executes in under a second and returns a fully-formed receipt with fee breakdown, seller protection status, and dispute categories.
## What's next for PayGuard AI
Near-term product improvements:
- Real merchant integration.Right now the payment goes to a sandbox merchant account. The next step is connecting PayGuard to a real product catalog or a commerce API (Shopify, Channel3, or similar) so the proposed transaction can be sourced from live inventory rather than typed in manually.
- Multiple policies per user.Currently one policy governs one transaction. The natural extension is a wallet of named policies — "Household", "Personal","Subscriptions" — each with its own budget and restrictions, and an agent that picks the right one.
- Refund and dispute visibility.The PayPal API already exposes refund links in the capture response. Surfacing those in the audit trail would make the tool useful after the transaction, not just before it.
Medium-term:
- A policy DSL.Let advanced users write policies by hand in a small, safe language rather than describing them in English. This would allow version control and review of spending rules, which a serious financial tool would need.
- Team accounts. Multiple humans authorizing from a shared budget, with the audit trail serving as the authorization record.
- Prompt-injection defense.The current architecture already prevents the model from *bypassing* a policy.The next layer is detecting attempts — a product description that tries to instruct the agent,for example — and flagging them to the user.
Long-term:
-The core idea generalizes.Any AI agent that has been granted a limited authority — spend up to $X,hire for up to Y hours,book travel within Z constraints — needs the same primitive: a policy authored by a human, enforced by code, that the agent cannot rewrite. PayGuard AI is one implementation of that primitive for payments. The pattern extends to any delegated authority an autonomous system is given.
The project is live at:
https://payguard-ai-qloi.onrender.com and the source is at:
https://github.com/Sule-Bashir/payguard-ai
under the MIT license.
Built With
- agentic
- agents
- ai
- api
- commerce
- css
- express.js
- fintech
- github
- groq
- html
- javascript
- llm
- node.js
- openai/gpt-oss-120b
- orders
- paypal
- render
- sandbox
- termux
Log in or sign up for Devpost to join the conversation.