Inspiration

AI agents can read emails, issues, web pages, and documents—and then act with powerful tools. That creates a dangerous gap: agents often cannot reliably distinguish information from instructions. A hidden prompt inside an email can convince an otherwise helpful agent to leak data, execute a command, or take an action the user never requested.

Most prompt filters ask whether text looks malicious. We wanted to ask a more useful question at the moment that matters: does this concrete action match the user's original intent?

What it does

Bouncer is an intent firewall that sits between an AI agent and its tools. Every proposed tool call is intercepted before execution, normalized into a concrete effect, and checked against the user's original goal.

Bouncer returns one of three decisions:

  • ALLOW — the action clearly serves the user's request.
  • ASK — the action is genuinely ambiguous and requires confirmation.
  • BLOCK — the action exceeds the user's intent and is never forwarded.

The result is context-sensitive security. Reading an inbox may be legitimate for “summarize my emails,” while forwarding that inbox to an unknown address is blocked—even if a malicious instruction inside an email told the agent to do it. Legitimate work can still continue after the attack is stopped.

How we built it

Bouncer is implemented as a TypeScript Model Context Protocol proxy. It captures the original user goal, discovers downstream MCP tools, intercepts each call, and normalizes it into READ, SEND, or EXECUTE semantics. It then sends a tightly scoped judgment request containing the goal, proposed effect, and bounded recent context to NVIDIA Nemotron Super through NVIDIA's hosted OpenAI-compatible API.

The enforcement layer fails closed: missing goals, invalid model output, audit failures, or uncertain dangerous actions never silently reach the downstream tool. Deterministic invariants block crisp risks such as secret-bearing outbound sends and execution of untrusted content, while Nemotron handles the semantic intent judgment that static rules cannot.

We also built:

  • A real MCP proxy and mocked email/GitHub tool servers for end-to-end demonstrations.
  • Bouncer Live, an animated visualization that shows tool calls approaching “the door” and being allowed or blocked.
  • A Python evaluation harness with deterministic, text-rule, Nemotron, and hybrid systems.
  • Public AgentDojo replay support and reproducible reports, dashboards, and demo artifacts.
  • A fully reproducible launch trailer with a Kokoro-generated voice-over.

Challenges we ran into

The hardest challenge was separating strong security claims from honest evidence. An early evaluation accidentally exposed risk labels to the evaluator and produced a suspiciously perfect result. We removed that leakage, reran the benchmark, and preserved the earlier run for transparency.

We also had to make the system robust to model failures. Hosted inference can time out, return malformed JSON, or produce a verdict whose explanation contradicts it. We added strict schema validation, bounded retries, consistency checks, secret redaction, and fail-closed behavior.

Finally, intent is not the same as keyword matching. A legitimate high-impact action may be implied by a user's goal without repeating its exact words. That is where Nemotron's semantic reasoning mattered most, but it also required careful prompts and bounded context to avoid trusting injected content.

Accomplishments that we're proud of

  • Built an enforceable tool boundary, not just another chatbot or text classifier.
  • Prevented all attacks in our leak-fixed 48-case diagnostic set while preserving most benign actions.
  • Completed an end-to-end 12-episode comparison where Nemotron stopped the paraphrased shell attack missed by text-only rules.
  • Made every major artifact reproducible: tests, evaluation reports, dashboard, live visualization, demo, and trailer.
  • Kept the threat model and limitations explicit instead of presenting a small benchmark as proof of universal security.

What we learned

Security for agents is fundamentally about authorization at the action boundary. Content can be adversarial, but harm occurs when an agent turns that content into a consequential tool call. Comparing the proposed side effect to the user's actual goal is more useful than trying to classify every piece of text as good or bad.

We also learned that model reliability is part of the security boundary. Structured outputs, redaction, audit logs, deterministic invariants, and safe failure modes are as important as raw model quality.

What's next for Bouncer

Next we want to add stateful trajectory reasoning, expand beyond READ/SEND/EXECUTE to WRITE, TRANSACT, and AUTHORIZE, run larger public and adaptive attack benchmarks, compare directly with additional guardrail systems, and add a user approval/resume flow for ASK decisions. We also want to package Bouncer as a drop-in policy layer for more MCP hosts and tool servers.

Links

Built With

Share this project:

Updates

Submission history