Inspiration

My mentor, a teacher, pointed out that no school hands a student the master keyring. They sign a hall pass: one student, one destination, expires soon. We're doing the opposite with AI agents. We hand them a live API key and a hopeful prompt. So we built the hall pass.

What it does

Warrant7 sits between an agent and the tool it wants to use.

  • The agent holds no credential. It submits an exact action.
  • Deterministic policy outside the model returns ALLOW, REVIEW, or BLOCK. Refunds up to $500 pass, $500 to $5,000 need a human, over $5,000 never execute.
  • On ALLOW we mint a one-use pass: bound to the request digest, one 256-bit nonce consumed atomically, dead in 120 seconds. Change the amount and the pass is void.
  • On REVIEW a human gets one sentence on their phone: what it costs, whether it's reversible, and a countdown. Silence is a denial.
  • Only the executor holds the real key, and it refuses anything without a valid, unspent pass.
  • Every event becomes a signed safety receipt: RFC 8785 canonical JSON, SHA-256 chained to the previous record, Ed25519 signed. Anyone can verify with just the public key.

How we built it

Next.js 15, TypeScript, Tailwind, Supabase Postgres and Realtime, Node crypto for SHA-256 and Ed25519, deployed on Vercel. Two more surfaces on the same backend: a native SwiftUI approver app wired through the /api/v1 gateway, and an MCP endpoint at /api/mcp with Bearer agent tokens, registered live in Claude Code. Any MCP-speaking agent inherits policy, approval, and receipts without changing a line of its own code.

Challenges we ran into

Canonicalization. Two identical objects hash differently if the keys are ordered differently, so verification failed for no real reason until we implemented RFC 8785 properly. Nonce consumption had to be atomic or a replayed pass would work twice. And the whole system had to fail closed, meaning if we can't reach a decision, nothing executes.

Accomplishments that we're proud of

We ran a real prompt injection against a live model. It took the bait and submitted the $2,400 refund. The money never moved. The model was manipulated, the safety layer wasn't.

Also: the chain verified end to end, and everything is green. 78 gateway tests, 55 in WarrantKit, 60 on iOS.

What we learned

Guardrails inside a prompt are a suggestion. Guardrails outside the model are a rule. And "trust us, here are the logs" isn't evidence. Evidence is something a regulator, an auditor, or an angry customer can check without asking us for anything.

What's next for Warrant7

The engine doesn't know what a refund is. It reads an action and a number. Next executors: outbound email to external domains, permission grants and password resets, prod database writes, and deploys. Then customer-authored rules and per-tenant signing keys.

Built With

  • api
  • ios
  • langfuse
  • next.js
  • openai
  • supabase
  • swfit
  • vercel
  • xcode
Share this project:

Updates

Submission history