-
-
A real AI agent was prompt-injected into paying an attacker 500 USDC. Leash's smart contract blocked it onchain.
-
LeashVault is deployed and source-verified on Base Sepolia (Exact Match).
-
Live vault stats and activity feed: 13 attack attempts blocked and recorded onchain, with the vault still active.
-
Reputation gating: the attacker scores 12/100, below the minimum of 60, so every payment to it is blocked.
-
The policy lives onchain: per-payment cap, daily cap, approval threshold, minimum reputation, and approval window.
-
Leash dashboard: every payment the AI agent makes is checked onchain, with real verified transactions linked below.
-
Owner controls: a one-click kill switch, human approval for large payments, and spending rules enforced by the contract.
-
Open-source contract built on OpenZeppelin's SafeERC20, Ownable, and ReentrancyGuard.
Inspiration
AI agents can now pay for things on their own. x402 lets them pay for APIs over HTTP, and ERC-8004 gives them onchain identity. But an agent that holds a private key has unlimited control over its funds, and LLMs can be prompt-injected by any webpage or API response they read.
Humans already lose huge sums by signing things they don't understand: wallet-drainer phishing took roughly $494M in 2024 (Scam Sniffer). Autonomous agents will make the same mistake at machine speed, with nobody watching.
Problem Statement
There is no onchain way to give an AI agent spending power without also giving it the power to lose everything.
- Guardrails written in prompts can be bypassed by prompt injection.
- Identity standards say who an agent is, not whether a payment is safe.
- A private key is all-or-nothing: there are no limits, no approvals, and no emergency stop.
What it does
Leash is an onchain firewall wallet for AI agents. The agent never holds the funds: it can only spend through the LeashVault smart contract, and every payment is checked onchain against rules the owner sets.
| Check | Rule | Reason code |
|---|---|---|
| Kill switch | Owner can freeze all spending instantly | VAULT_PAUSED |
| Validity | No zero amounts or invalid recipients | INVALID_PAYMENT |
| Per-payment cap | Max 10 USDC per payment | EXCEEDS_PER_TX_CAP |
| Daily cap | Max 25 USDC per day | EXCEEDS_DAILY_CAP |
| Reputation | Unknown recipients need a score of 60+ (allowlist bypasses) | LOW_REPUTATION |
| Balance | Vault must hold enough funds | INSUFFICIENT_BALANCE |
| Human approval | Payments above 5 USDC wait for the owner, with expiry | pending |
Blocked attempts don't just fail: they're recorded onchain as PaymentBlocked events with a reason, so every attack leaves permanent evidence.
Solution & Unique Value
You can't make an LLM immune to prompt injection, but you can make the money immune.
- Enforcement lives in the contract, not the prompt. Even a fully hijacked agent physically cannot overspend or pay an untrusted address.
- Attacks become evidence onchain.
- Human-in-the-loop approval for large payments, and a kill switch.
- x402-style payments settled through a policy vault: the API answers
402 Payment Required, the agent pays through LeashVault, and the server verifies the receipt onchain with route-bound, single-use nonces. - Reputation gating designed to plug into ERC-8004.
🔥 A real AI got hijacked, and Leash stopped it
During testing, a real LLM agent (gpt-oss-120b) read a malicious API response with a hidden "system notice" telling it to pay 500 USDC to an attacker. The model obeyed, and the vault blocked it onchain. No funds moved. 👉 See the blocked attack on Basescan
A fully hijacked agent that then tried smaller amounts (4 and 2 USDC) to slip under the cap was blocked too, because the attacker's reputation score (12) is below the minimum (60).
How we built it
- Smart contracts: Solidity + Foundry + OpenZeppelin (
SafeERC20,ReentrancyGuard,Ownable). 42 tests, including fuzzing, which proves the daily cap can never be exceeded. Deployed and verified on Base Sepolia. - Paywalled API: Node.js + Express + viem. It issues x402-style
402responses and verifies payments by decoding vault events from the transaction receipt, with replay protection. It includes simulated malicious endpoints for testing. - AI agent: TypeScript with LLM tool calling (OpenAI-compatible / Anthropic), a scripted mode, and a simulated hijacked-agent mode. It streams every step over Server-Sent Events.
- Dashboard: Next.js + Tailwind + wagmi + viem. It shows live vault stats, a real-time onchain activity feed, "Attack blocked" alerts, pending approvals, a kill switch, and a live agent console, with no mock data. It's deployed on Vercel.
- CI: GitHub Actions runs
forge test, type-checks, and builds on every push.
Technology Stack
- Blockchain: Base Sepolia (EVM)
- Smart contracts: Solidity, Foundry, OpenZeppelin
- Protocols: x402-style HTTP 402 payments, ERC-20, ERC-8004-ready reputation interface
- Backend: Node.js, Express, TypeScript, viem
- AI: LLM tool calling (gpt-oss-120b via Groq)
- Frontend: Next.js, React, Tailwind CSS, wagmi
- Infra: Vercel, GitHub Actions
Challenges we ran into
- Recording blocked attempts without reverting. A revert erases events, so the vault emits
PaymentBlockedand returns instead. That's what makes attacks visible onchain. - Approval expiry. Reverting would also undo "mark as expired", so expired approvals return
falseand emit an event instead. - Payment replay and route confusion. Nonces are route-bound and single-use, so a 0.01 USDC weather payment can't unlock an 8 USDC report.
- Public RPC limits. The dashboard loads event history in adaptive 1,000-block chunks.
- Proving the danger is real. Modern models sometimes catch injections, so we built both a real-LLM scenario and a deterministic hijacked-agent simulation. In our tests, the real model did fall for it.
Accomplishments that we're proud of
- A real LLM was prompt-injected, and our vault stopped it live on a public blockchain.
- A complete end-to-end system (smart contracts, payment protocol, AI agent, and dashboard) working on testnet today.
- 42 passing tests, verified contracts, and CI on every push.
What we learned
- Security for AI agents can't live only inside the model. The safest place for spending rules is the one place the agent can't talk its way around: a smart contract.
- Designing contracts for observability (recording failures, not just successes) turns a security tool into an audit trail.
- The x402 pattern composes cleanly with onchain policy: the HTTP layer asks for payment, and the chain decides whether it's allowed.
What's next for Leash: Onchain Firewall Wallet for AI Agents
- ERC-8004 reputation adapter to replace our mock oracle with real agent reputation
- ERC-4337 / EIP-7702 smart-account module so any agent wallet can use Leash
- Per-merchant budgets and time-based spending rules
- Multi-chain support and an SDK for agent frameworks
- A security audit, then mainnet with real USDC
Links: 🌐 Live demo · 💻 GitHub · 📜 Verified vault · 🎥 Demo video · 📊 Pitch deck
Testnet only: mock USDC and mock reputation oracle. Not audited.
Built With
- base
- erc-8004
- ethereum
- express.js
- foundry
- github-actions
- next.js
- node.js
- openzeppelin
- react
- smart-contracts
- solidity
- tailwindcss
- typescript
- vercel
- viem
- wagmi
- x402
Log in or sign up for Devpost to join the conversation.