Inspiration
My research on prompt injection showed how easily AI agents act on whatever an attacker puts in front of them. Once agents can pay invoices, that turns into fraud. The oldest scam in accounts payable, "we changed our bank details," works even better on an agent than on a person. I wanted to make agent payments safe to turn on.
What it does
Countersign sits between an AI accounts-payable agent and the money. The agent reads invoices and proposes payments but never holds a key. Countersign checks each payment: where the wallet address came from, hidden text, lookalike domains, and the vendor's payment history. It then pays instantly, asks the owner to approve on their phone, or blocks. A live dashboard shows every decision and why.
How we built it
- App and agent: a Next.js web app with an agent built on Claude and the Vercel AI SDK.
- Risk engine: fixed rules that score each payment and decide to pay, ask a human, or block. A separate Claude call with no tools flags suspicious emails as one small input.
- Approvals: Auth0 handles login and push approvals to your phone through CIBA and the Auth0 Guardian app.
- Risk checks: Tiger Data stores payment history and runs the time-based checks: spending spikes, unusual amounts, and duplicate invoices. It also holds the agent's event log, which powers the live dashboard and replays.
- Payments: they go out on Solana devnet as mock-USDC transfers. Each carries an on-chain memo recording the decision and who approved it.
- Testing: an eval harness runs every attack with the guard off and on.
Built in part with Claude Code.
Challenges we ran into
- Phone approvals took the most setup. CIBA needs the right grant type, a notification channel, and a verified approver enrolled in Guardian. The approval message is capped at 64 plain characters, so each payment's amount and wallet had to fit in that space.
- Keeping the attack honest. I wanted the unguarded agent to fall for the scam without being rigged to. It uses a realistic accounts-payable prompt and the same tools in both modes, so the scam emails had to be convincing enough to fool a capable model.
- Not making the agent wait on a human. While a payment waits for approval on the phone, the agent keeps working through the inbox. The app checks back with Auth0 at the pace Auth0 asks for.
Accomplishments that we're proud of
- The same agent, prompt, and inbox produce opposite outcomes. With the guard off, it pays the attacker in a real Solana devnet transaction. With it on, the attack is stopped.
- Security comes from the architecture, not the prompt. The agent never holds a key, and the gateway, not the model, decides when to ask a human. The signer refuses any payment without a clean decision or a verified approval.
- Every decision explains itself in plain English. The dashboard can trace a stolen wallet address back to the email it came from.
- The approval on the phone shows the actual payment, not a generic "approve?" prompt.
What we learned
- The model should never decide when it needs permission, because a prompt can't be trusted to enforce a rule.
- Where a wallet address came from is a stronger fraud signal than anything the email says.
- Human approval only works if people see the real details and aren't asked too often. That's why normal invoices pay without interrupting anyone.
- Payment risk is a time-series problem. History is what separates a legitimate bulk order from fraud.
What's next for Countersign
- Package the gateway as middleware or an MCP server that sits in front of any agent's payment tool.
- Support real payment rails like Stripe, ACH, and USDC, since the guard doesn't depend on Solana.
- Add out-of-band vendor verification, so a changed bank detail is confirmed with a call to a known number before it's ever used.
- Enforce spending limits on-chain with a Solana program, and expand the eval to more models and attack types.
Built With
- anthropic
- auth0
- claude
- next.js
- postgresql
- react
- solana
- tailwind
- timescaledb
- typescript
- vercel
Log in or sign up for Devpost to join the conversation.