Inspiration
AI agents are getting good at acting autonomously. They browse the web, call APIs, write and execute code, and interact with third-party services on behalf of users. But this autonomy creates a massive blind spot. If an agent gets tricked by a prompt injection hidden inside a RAG document, or if it accidentally leaks an API key in an outbound request, there's nothing sitting between the agent and the open internet to stop it.
Existing security tools weren't built for this. API gateways protect servers from malicious clients, but with AI agents the threat model is flipped. The agent is the client, and it can be manipulated into doing harmful things without even knowing it.
We kept thinking about how network firewalls changed everything in the 1990s. Firewalls didn't trust applications to behave correctly. They enforced policy at the network boundary. We wanted to build the same thing for AI agents. What if every single HTTP request an agent makes had to pass through a security checkpoint, transparently, without requiring any code changes?
What it does
Bastion AI is a security gateway that sits between AI agents and the outside world. Agents run inside isolated Docker containers where all their network traffic gets routed through a TLS-intercepting proxy. A dedicated Guard service analyzes every request and response in real time using a multi-layered detection pipeline.
The prompt injection detector combines 19 regex patterns covering instruction override, identity hijack, jailbreaks, and encoding tricks with a fine-tuned transformer classifier. We built the classifier by fine-tuning Meta's Llama Prompt Guard 2-86M on a mix of four public datasets and our own synthetic samples generated by Claude. Scores from both layers are fused conservatively using $\text{score} = \max(s_{\text{regex}}, s_{\text{model}})$ so that either detector can independently flag a threat.
RAG content sanitization scans documents fetched from retrieval systems for embedded instructions, hidden markers, and exfiltration attempts. Content scoring above 0.5 gets sanitized in-flight (HTML stripping, URL removal, Unicode normalization, zero-width character removal). Content above 0.8 is quarantined and never reaches the agent.
Secret and PII detection catches API keys, cloud credentials, passwords, and personally identifiable information before they leave the sandbox.
Egress policy enforcement uses a domain allowlist to restrict which external services agents can contact, preventing unauthorized data exfiltration at the network level.
Full audit trail logs every intercepted request with the decision rationale, risk score, threat classification, and latency. Everything is queryable through a real-time monitoring dashboard that shows KPIs like threats blocked, secrets redacted, and RAG chunks quarantined, along with decision distribution charts and threat timelines.
How we built it
The system has four layers stacked on top of each other.
Agent Sandbox (isolated Docker container)
| ALL outbound HTTPS traffic
v
mitmproxy Gateway (TLS interception)
| Decrypted request/response payloads
v
Guard Service (FastAPI security engine)
| ALLOW / DENY / SANITIZE decision
v
Upstream Services (OpenAI, Anthropic, tool APIs, etc.)
The proxy layer is a custom mitmproxy addon that hooks into every HTTP request and response. It auto-detects route types by matching against known LLM API hosts (OpenAI, Anthropic, Google, Mistral, Cohere) and supports an override header for custom routing.
The Guard service is a FastAPI application that runs the full scanning pipeline. Each request passes through injection detection, policy checks, RAG scanning, and 30+ configurable llm-guard scanners. We keep decisions under 500ms by warming up all ML models at startup and handling audit logging asynchronously.
Agents run in minimal Docker containers with HTTP_PROXY and HTTPS_PROXY environment variables pointing to the gateway. A shared Docker volume distributes the mitmproxy CA certificate, and a boot script installs it into the system trust store before launching the agent. This makes interception completely transparent to Python requests, httpx, curl, and Node.js without touching a single line of agent code.
The dashboard is a React + TypeScript SPA built with Vite, TailwindCSS, and Shadcn/ui. It pulls real-time metrics from the Guard's admin API using TanStack React Query and renders charts with Recharts.
We also built a CLI tool called agw that lets developers scaffold agent projects, run them through the gateway, or execute ad-hoc Python code inside the sandbox.
Network isolation is enforced through three Docker bridge networks. The workload network connects agents to the proxy. The control network connects the proxy to the Guard. The external network gives the Guard access to the internet. Agents physically cannot bypass the proxy because there's no network route available to them.
Training the injection classifier
The prompt injection classifier is a fine-tuned version of Meta's Llama Prompt Guard 2-86M, an 86M-parameter mDeBERTa model already pre-trained for injection detection. We improved it by training on a combination of four public HuggingFace datasets (safe-guard, SPML, deepset, Microsoft llmail-inject-challenge) plus synthetic data we generated ourselves.
For the synthetic data, we built a parallel generation scaffold that spawns seven Claude workers simultaneously, one per attack category (override/hijack, exfiltration, tool hijack, policy bypass, mixed attacks, benign samples, and hard negatives). Each worker calls Claude in a loop with category-specific prompts and a handful of randomly sampled seed examples from the public datasets as style references. Claude returns structured JSON matching a strict schema, which gets validated, deduplicated by SHA256 hash, and written to a SQLite database. Hard negatives are particularly important here. These are samples that contain prompt-injection terminology (like "ignore previous instructions" appearing in an educational discussion) but are actually benign. Without them, the classifier learns to trigger on keywords instead of intent.
We fine-tuned with PyTorch Lightning using weighted cross-entropy to handle class imbalance, a cosine learning rate schedule with warmup, and early stopping on validation F1. The embedding layer is frozen to save memory since the multilingual embeddings add ~190M parameters on top of the 86M transformer. The exported model gets dropped into the guard service and runs inference at request time with no external API calls needed.
Challenges we ran into
Getting TLS interception to work reliably across every HTTP client library was harder than expected. Python, Node.js, and curl each have their own way of handling certificate trust. We had to orchestrate multiple environment variables (REQUESTS_CA_BUNDLE, SSL_CERT_FILE, NODE_EXTRA_CA_CERTS, CURL_CA_BUNDLE) and write an init script that blocks until the certificate volume is populated before the agent starts.
Latency was a constant concern. Every agent request passes through two additional network hops, and running ML classifiers synchronously risked timeouts. We solved this by pre-loading models during startup, logging to the database asynchronously via FastAPI BackgroundTasks, and sampling low-risk traffic (only 2% of allowed tool calls get logged, while every threat event is captured at 100%).
Combining regex and model-based injection scores without drowning in false positives took a lot of tuning. We landed on conservative max-fusion where $\max(s_{\text{regex}}, s_{\text{model}})$ determines the final score. For a security system, it's better to over-flag than to let an attack slip through.
Generating high-quality synthetic training data was its own challenge. Early runs produced samples that were too formulaic. The classifier would memorize surface patterns like "ignore previous instructions" instead of learning to detect the underlying intent. We had to carefully design hard-negative prompts that use injection-related language in benign contexts (security tutorials, academic discussions, quoted examples) to teach the model the difference between talking about prompt injection and actually attempting it.
Accomplishments that we're proud of
Agents don't need any SDK, library, or code modification. Just run inside the sandbox container and all traffic is automatically protected. This means Bastion works with any language, any framework, and any LLM provider out of the box.
The SANITIZE decision doesn't just flag suspicious content. It actually rewrites the HTTP request body in transit, stripping malicious payloads before they ever reach the upstream service.
We implemented split-token authentication where proxy and admin scopes use separate tokens compared with constant-time HMAC. Even if an agent container is fully compromised, the attacker can't access the admin dashboard or modify security policies.
We built an end-to-end ML pipeline that goes from public datasets and Claude-generated synthetic samples all the way to a deployed fine-tuned classifier running inference inside the guard service. The whole loop, from data generation to model export to live detection, is reproducible and self-contained.
The system covers a wide threat surface from a single enforcement point. Prompt injection, RAG poisoning, secret exfiltration, jailbreaks, encoding attacks, PII leakage, and unauthorized tool calls are all caught at the network boundary.
What we learned
Operating at the network layer is fundamentally more robust than SDK-based security. By sitting at the proxy, we protect every agent regardless of its implementation, without asking developers to adopt a new library or change their code.
Regex and ML classifiers are better together than either one alone. Regex patterns are fast, interpretable, and reliable for known attack signatures. ML classifiers catch novel attacks that look semantically similar to known patterns but use different phrasing. Max-fusion gives the best of both worlds.
Using one LLM (Claude) to generate adversarial training data for a smaller specialized model is a surprisingly effective workflow. Claude can produce diverse, realistic attack samples across categories that would take a human researcher days to write by hand. The key insight is that the synthetic data is only as good as the prompts and the hard negatives. Without deliberate effort to include benign samples that look like attacks, the classifier just learns keyword matching.
Agent security has to be bidirectional. It's not just about what agents send outbound. Responses from upstream services can also contain prompt injections designed to hijack the agent's behavior. Scanning both directions turned out to be essential.
Observability and security are the same thing in practice. The audit trail and dashboard aren't extras. They're how you discover new attack patterns, tune detection thresholds, and build confidence that your policies are actually working.
What's next for Bastion AI
We want to add anomaly detection that uses historical traffic patterns to flag behavioral outliers, like an agent suddenly contacting an unusual domain or sending abnormally large payloads.
Multi-tenant deployment is on the roadmap so multiple teams can share a single gateway with per-tenant policies, quotas, and audit isolation.
We're planning policy-as-code support where egress rules, scanner thresholds, and response actions are defined in YAML files that can be version-controlled and reviewed in CI/CD.
Real-time alerting through webhooks to Slack, PagerDuty, and email will let teams respond immediately when high-severity threats are detected.
Finally, we want to close the feedback loop by piping confirmed false positives and negatives from the dashboard back into classifier retraining, so the system gets smarter over time.
Built With
- docker
- docker-compose
- fastapi
- framer-motion
- hugging-face-transformers
- javascript
- llm-guard
- mitmproxy
- python
- pytorch
- radix-ui
- react
- recharts
- shadcn/ui
- spacy
- sqlalchemy
- sqlite
- tailwindcss
- tanstack-react-query
- typescript
- vite
- vitest
- zod
Log in or sign up for Devpost to join the conversation.