Inspiration
AI agents are starting to do real stuff: read your inbox, send emails, delete files, even move money. That's exciting, but it also freaked us out a little. The moment an agent can act, it's also reading untrusted content all day, and that content can hide instructions that hijack it.
The scarier part? When the agent asks you to approve something, that confirmation shows up on the same screen the software controls. A hijacked agent can describe wiring money to a fraudster as "paying your vendor," and you'd hit Approve on something you never actually saw.
Then we realized this problem has already been solved, twice. Crypto has hardware wallets, and banks have separate confirmation devices like chipTAN for online transfers. Both show you the real transaction on a screen the computer can't touch. So we asked ourselves: why don't AI agents get one of those? That question turned into Gatekeeper.
What it does
Gatekeeper is a little hardware approval device (an ESP32 with its own screen) that sits between an AI agent and anything consequential. The agent can only ask. Nothing happens until the device shows the real recipient, file or amount on its own display and signs the verdict with an Ed25519 key that never leaves the hardware.
- It shows the truth. The OLED displays the actual action, including a
TAINTEDflag when the session has read outside email, and the real payee hiding behind a lookalike name. - It needs a human. You hold a button for 2 seconds to approve, and payments over \$500 also need an RFID card tap.
- It catches lies. If the agent claims an action is "low risk" when it isn't, the device locks the whole session.
- The bank checks, not the laptop. For payments, a separate process (standing in for a bank) re-checks the device's signature before any money moves through the Capital One Nessie sandbox. So even a fully compromised laptop can't forge an approval.
- It talks. Gatekeeper announces the real action out loud with ElevenLabs, like "Blocked: payment of \$750 to an unknown account."
How we built it
- Firmware (ESP32, PlatformIO/Arduino): request checks, the policy, the OLED, buttons, an RGB LED, an MFRC522 RFID reader, and Ed25519 signing with Monocypher. The signing key is generated on the device's first boot and stored in its flash.
- Host (Python, mostly standard library): an executor that holds no keys and verifies signatures against a pinned device key, an LLM agent loop, and a mock device that reproduces the firmware's verdicts byte for byte, so we could test without the board.
- Payments: a Nessie client (accounts, transfers, merchants, purchases) and a standalone bank process that verifies the device's signature before paying anyone.
- Benchmark: 24 attack scenarios and 20 normal ones, run through the real agent, executor and device.
- Demo: a click-through web dashboard, a live weak-model hijack, and an interactive "Beat the Gatekeeper" challenge.
The whole security story comes down to one signed message. The device signs
$$\text{msg} = \texttt{v2|act|to|file|fh|bh|amt|nonce|exp|taint}$$
and the bank only acts when
$$\text{verify}_{pk}(\text{msg}, \sigma) = \text{accept}$$
under the pinned device key, with a fresh nonce and a timestamp that hasn't expired. The laptop never holds the private key, so it simply can't produce $\sigma$.
Challenges we ran into
- Byte-for-byte agreement. The firmware and the Python mock had to build exactly the same signed string, or nothing would verify. One stray character and everything breaks, so this ended up shaping our whole contract design.
- A real security bug. On macOS,
data/SENSITIVE/tax_return.pdfopens the same file as the lowercase path, and it slipped right past our "sensitive" check. We made the match case-insensitive. - A gap our own benchmark caught. One attack, mass-deleting public files, got through because ordinary deletes were auto-approved. We closed it: a tainted session can no longer auto-approve deletes.
- Hardware weirdness. Our USB-serial adapter rebooted the ESP32 every time we connected, wiping its memory between calls. That made the session lock really confusing to debug until we figured out what was going on.
- The Nessie sandbox. Its Bills endpoint is broken (you can create a bill but can't read it back), and it can't delete transactions. So we modeled vendor payments with Merchants and Purchases instead, and reset the demo at the display layer.
What we learned
- Detection is a guess; enforcement isn't. We tested real models against our phishing invoice. Frontier models like Grok 4 resisted it, but DeepSeek, Llama 3.3 and Ministral-8B all fell for it instantly, and one even lied in its narration about what it had done. Lesson learned: you can't bet money or data on the model being careful. The device has to be the backstop.
- The hard part is the threat model, not the crypto. Deciding exactly what we trust (the device, its key) and what we assume is compromised (the laptop, the agent) is what makes the design hold up. It's also our answer whenever someone asks "why hardware?"
- Being honest about limits makes the project stronger. Saying clearly what we don't protect against made Gatekeeper more convincing, not less.
Accomplishments that we're proud of
- Real Ed25519 verdicts coming off a \$30 microcontroller, verified end to end.
- 24/24 attacks caught, 0 executed, 0 false positives, about 42 ms decision time on the real board.
- A current model (DeepSeek) getting hijacked live on stage, and the device stopping it.
- A forged, tampered or replayed approval rejected by the bank, live. Plus RFID card co-signing and live revocation of a lost card.
What's next for Gatekeeper
A dedicated secure chip for the key, a production-grade bank and mail-gateway verifier, per-organization policies, and taking the same device to AI coding agents, so it can gate destructive commands and secret leaks, not just payments.
Built with
ESP32, PlatformIO, Arduino, Monocypher and PyNaCl (Ed25519), Python, pyserial, Capital One Nessie, ElevenLabs, and an OpenAI-style LLM endpoint.
Log in or sign up for Devpost to join the conversation.