Inspiration
Everyone is racing to give AI agents real power. They can send email, move money, run shell commands, and call APIs on their own now. Almost nobody is asking the obvious follow-up question: what actually stops an agent when it gets something wrong? Input guardrails check text before it reaches the model. Tracing tools tell you what happened after the fact. Nothing sits at the exact moment an agent tries to do something and says no. We wanted to build that missing piece. A firewall for AI agents that looks at every action an agent tries to take and decides, right there, whether to allow it, block it, or hand it to a human.
What it does
AgentGate sits between an AI agent and its tools as an MCP proxy. Every tool call the agent makes passes through it first and comes back as allow, block, or escalate, along with a written reason that names the exact policy that fired. All of this happens before the action runs, not after.
The obvious stuff gets caught instantly. A fast rule layer flags things like PII (with a real Luhn check on card numbers), destructive commands such as rm -rf and DROP TABLE, spending limits, and rate limits in about a millisecond, with no model call at all. Anything more subtle goes to a five-node LLM judge that reasons about the action against your written policies. On top of that, a pattern detector watches the whole session and catches things no single check ever could, like thirty separate \$400 purchases that quietly add up to \$12,000.
The part we like most is that policies are just plain-English markdown you drop in. You can hand it a text file, a PDF, or a Word doc, and it becomes an enforceable policy on the very next tool call. No redeploy, no code change.
How we built it
It is a monorepo with a few moving parts that each do one job.
The gateway is TypeScript. It is the MCP proxy that mirrors whatever tool server the agent already uses, so the agent never knows anything changed. It runs the five deterministic rules, writes every decision to a Cloudflare D1 audit log, reports errors to Sentry, and also ships as a Cloudflare Worker.
The engine is Python, built on FastAPI and LangGraph. It is a five-node graph: a classifier, a policy retriever, a risk judge, a decision gate, and the pattern detector. Every node emits a LangFuse span so we can see exactly where time goes. Retrieval is hybrid, mixing Cloudflare Vectorize embeddings with BM25 keyword search. The final decision is plain threshold logic rather than a guess: below 30 is allow, 70 and above is block, and anything in between escalates to a human.
Three mock agents (customer support, procurement, and a coding agent, with nine tools between them) run their calls through the gateway, so the demo shows the firewall working live instead of us just describing it. A Next.js dashboard shows the live feed, analytics, the policy editor, and a review queue. A TypeScript eval harness scores the whole system against 112 labelled scenarios so we always know where we stand.
Challenges we ran into
The hardest part was not building the judge. It was learning to trust our own scores.
Early on, our test scenarios assumed a \$10,000 spending limit, but the policy the engine actually shipped with used a lower one. So a lot of our "failures" were really just our labels disagreeing with the real policy. Fixing that honestly, instead of quietly relabeling things to make the number go up, turned into its own small engineering project. We built a flag that marks whether a run is an official product number and banners it loudly if it is not, a hash that refuses to compare two eval runs scored on different scenario sets, and a rule that a failing run parks itself in a separate file instead of overwriting our good baseline. That last one saved us more than once.
We also hit a real calibration bug and chose to understand it rather than paper over it. Our judge tends to score an over-budget but otherwise ordinary purchase high enough to block it, when the policy actually says it should be escalated to a human instead. So the policy text and the pipeline disagree, and that is exactly why our escalate recall is our weakest number. The engine gets about 72 percent of the 112 scenarios right overall, and almost every miss is the same shape: it blocks something the policy only wanted escalated. In other words, when it is wrong, it is wrong on the side of being too strict, which for a security tool is the failure you would pick.
Accomplishments that we're proud of
Block recall is 100 percent. Across all 112 scenarios, the engine never once allowed an action we had labelled dangerous.
One policy language covers completely different worlds. The same cumulative rule that catches someone splitting a big purchase into thirty small ones also catches data being siphoned out of a network, just by counting bytes instead of dollars. Adding a whole new kind of detection means writing a markdown file, not shipping code.
The Zip integration surprised us. With live budget grounding turned on, the exact same \$4,000 purchase order gets escalated when Zip is not consulted and blocked when it is, with the reason that it would push Marketing's Q3 budget over its limit and that the vendor is not approved. It even works the other way: a tiny \$180 order gets escalated because the budget is already 97 percent spent. No fixed threshold in any policy file could catch that.
Our held-out set is the number we are proudest of. We wrote 12 fresh scenarios straight from policy text, never from anything the engine had produced, to test policies we added later. The engine got 11 of 12 right, about 92 percent. That is the number that answers "did you just overfit to your own benchmark," and the answer is no.
And then there are the relabels. We corrected 12 scenario labels during the project, and every single correction moved a label toward escalate, which is the class the engine is worst at. We made our own score lower on purpose, because the labels should match the policy, not match whatever the model happens to output.
What we learned
We learned that the honest number is worth more than the flashy one. Our held-out 92 percent means more to us than any headline accuracy, because it proves we were not grading ourselves on a curve we drew ourselves.
We learned that a firewall has to fail closed. If the judge stops responding, our gateway refuses to start at all, because a security tool that quietly falls back to running half its checks is more dangerous than one that stops and tells you something is wrong.
And we learned why the placement matters so much. Because the agent is pointed at AgentGate instead of at its real tools, the agent never even holds the credentials for the upstream server. It physically cannot go around the firewall, because there is nothing else for it to call. That is the whole idea in one sentence, and it is what makes this different from guardrails that only read the model's text.
What's next?
The first thing is closing that calibration gap, so over-budget but ordinary actions get routed to a human instead of blocked. After that, a review queue that actually saves its decisions instead of holding them in memory, deploying the Python engine to a real host so it is not living on a laptop, and running our network log analysis mode against real datasets. Longer term, we want more policy dimensions. A hospital could count patient records, a SaaS company could count exported rows, and each one is just another markdown file.
Built With
typescript, python, mcp, langgraph, fastapi, next.js, react, tailwind, cloudflare-workers, cloudflare-d1, cloudflare-vectorize, openai, zip, sentry, langfuse, mongodb, docker, turborepo
Try it out
https://github.com/japneet250/Agent-Gate
(Runs locally with ./demo.sh, which boots the engine, gateway, and dashboard together with health checks. There is no public deployment yet.)
Submitted to
Zip, Cloudflare, Sentry, Warp, OpenAI
Created by
Japneet Singh (japneet250), Megh Joshi (Megh2k4), Aaryan Paiva (Aaryan-Paiva)
Built With
- claude
- cloudflare
- deepevals
- langfuse
- langgraph
- llmops
- mcp
- node.js
- openai
- python
- react
- sentry
- tailwind
- typescript
- zip
Log in or sign up for Devpost to join the conversation.