Inspiration
AI coding agents now run shell commands, edit files, and fetch URLs on your machine, and they sometimes trust whatever they read. A single poisoned README or web page can tell an agent to quietly send your .env or AWS keys to an attacker. Existing permission prompts don't help much: you either click "allow" on everything out of fatigue, or you read every command and lose the point of having an agent. We wanted something that understands why an action looks dangerous and can show what it would do before it does it.
What it does
Gatekeeper is a local server that sits between AI coding agents (Claude Code, VS Code Copilot, Codex) and your machine. Every action is normalized and given a verdict of allow, ask, deny, or question. Cheap rules settle most actions in milliseconds, so safe commands pass quickly, and never-allowed patterns are blocked outright. Riskier commands get a dry run in a throwaway Docker container with no internet access, fake "tripwire" secrets, and a logger that records every host the command tries to reach. Gatekeeper also tracks prompt injection by scanning web pages, MCP output, and package files for hidden instructions, so a risky action that follows an untrusted read can be traced back to its source. An LLM judge rates the risk and explains it in plain language, and every decision is logged with a git checkpoint so you can review it and roll it back. There is also a joke "useless mode" that denies destructive commands with a sarcastic question you have to answer.
How we built it
The core is a Python FastAPI server on localhost, backed by SQLite in WAL mode with numbered migrations and Pydantic types shared across every stage. Claude Code connects through four curl hooks, while VS Code and Codex connect as an HTTP MCP server that replaces their built-in terminal, file, and fetch tools, so Gatekeeper both decides and does the work. The decision pipeline parses commands with bashlex, applies a YAML rules file, and then runs the LLM judge and the Docker sandbox in parallel. The final verdict comes from a plain if/else step rather than an LLM, so a manipulated judge can't argue its way past a tripwire hit. The sandbox sits on an internal-only Docker network with a mitmproxy container that decrypts HTTPS and scans URLs, headers, and bodies for tripwire values in raw, URL, hex, and base64 forms. To measure all of this, we built an eval harness that runs attack and benign scenarios against a real Claude Code agent with the hooks on and off, plus a red-team loop where an attacker LLM tries to break it.
Challenges we ran into
The hardest part was connecting Gatekeeper to Copilot for live testing. VS Code and Codex have no hooks, so Gatekeeper can't simply watch them; it has to be their tools. That meant registering our server as an HTTP MCP server, turning off the built-in terminal, file edit, and fetch tools, and debugging real agent sessions end-to-end, including auth tokens, session tracking through MCP session IDs, and approval prompts that depend on whether the client supports elicitation. When a client can't show a prompt, we fall back to a separate gatekeeper approve terminal. Hand-testing with a live agent was slow, and every client behaved a little differently.
Other problems were quieter but just as real. Claude Code never tells hooks what you chose in its approval prompt, so the server infers approval when the matching tool_use_id shows up in the after-tool hook. We also had to make Gatekeeper fail-safe, so if the server is down, the hook denies instead of letting actions through. Sandboxing untrusted code safely was taken care of as well: repo images are built with --ignore-scripts, .devcontainer is never used, and saved changes from a run are treated as untrusted archives that can't escape the repo or touch .git.
Accomplishments that we're proud of
We're proudest of how rigorously we tested it. The test suite uses real captured Claude Code hook payloads and real MCP payloads instead of mocks, Docker-backed tests verify that the sandbox network truly has no route out, and CI runs lint and tests on every push.
We also measured Gatekeeper against a real Claude Code agent instead of just claiming it works. Across 10 scenarios, including poisoned READMEs, malicious npm postinstall scripts, curl piped into a shell, CI tampering after a web fetch, and exfiltration through environment variables, attacks succeeded 67% of the time (4 of 6) with the hooks off and only 17% of the time (1 of 6) with Gatekeeper on. We lead with success rate because a real agent sometimes refuses bad instructions by itself, so comparing the same attacks with and without Gatekeeper is the fairest measure. On speed, rules-only decisions take under 20 ms and judge plus sandbox takes 2 to 8 seconds, all through one pipeline that serves three different agent clients.
What we learned
We learned that deterministic checks should have the final say while LLMs only advise, which makes the system much harder to manipulate. Provenance also turned out to matter more than pattern-matching, because knowing what the agent just read is a stronger signal than how a command looks. Every agent client exposes different integration points, so the design has to be built around the weakest one. Measuring honestly changed how we saw the project as well: the eval showed us exactly where the rules were too noisy or too permissive, which is more useful than a perfect-looking demo. Finally, security tools only get used if they're fast and quiet, so most of our effort went into making the safe path better.
What's next for Gatekeeper
With the contributor-docs attack from our first eval now fixed, the next step is bringing MCP support to VS Code, so Gatekeeper can guard Copilot agent mode with the same decision pipeline it already uses for Claude Code. After that, we want to grow the eval well beyond 10 scenarios, with more repeats per scenario, multi-step attacks, real external hosts instead of a local canary, and Codex coverage. We also plan to support more agents, make gatekeeper init a one-step setup, add per-repo policy sharing for teams, and build a visual review dashboard on top of gatekeeper logs and rollback.
Log in or sign up for Devpost to join the conversation.