💡 Inspiration

Over the last year, software engineering crossed an inflection point: we moved from chatbots to Autonomous AI Coding Agents (Cursor, Claude Code, Windsurf, Devin, GitHub Copilot Workspace). These agents are granted direct terminal execution, filesystem access, and package manager control on developers' host machines.

However, agents blindly ingest untrusted codebases into their reasoning context. When inspecting open-source repositories or third-party Pull Requests, an attacker can embed disguised natural language directives in code comments, markdown docs, or environment configs. The agent reads the hostile prompt, follows it, and executes arbitrary shell commands on the developer's laptop—granting attackers full Remote Code Execution (RCE) and exfiltrating AWS tokens or SSH keys.

Existing security tools like CodeRabbit, Snyk, and Semgrep scan human code for syntax bugs after the fact. No tool existed to protect the AI agent itself from being hijacked before it executes code. That inspired us to build AgentGuard.


🛡️ What It Does

AgentGuard is an active, pre-execution security firewall and attack kill-chain correlator designed specifically for autonomous agentic IDEs.

Instead of acting as a passive PR linter, AgentGuard provides runtime defense:

  • Active Pre-Execution Firewall (guard): Wraps tool and shell commands. If an agent attempts to touch a poisoned file or run a smuggled CLI argument (e.g. -X_upload_payload.txt, --checkpoint-action=exec=...), AgentGuard hard-blocks the command at the OS level with Exit Code 126 before it touches the host system.
  • Differential Intent Analysis via NVIDIA Nemotron 3 Ultra 550B: Distinguishes between benign developer guidelines (e.g., project milestones, coding styles in .cursorrules or .mdc files) vs. malicious exploits (disabling sandboxes, credential exfiltration), achieving zero false alarms.
  • Multi-Stage Attack Kill-Chain Correlator: Links isolated findings into complete narrative graphs (CHAIN-RCE-001, CHAIN-EXFIL-002) showing the entire multi-stage flow: Ingress ➔ Persistence ➔ Action.
  • Cryptographic Baseline Trust (--bless): Signs legitimate developer guidelines into .agentguard.lock with SHA-256. If an attacker tampers with an approved rule in a Pull Request, AgentGuard raises an instant Critical Baseline Integrity Violation.
  • Autonomous Rule Discovery (--update-rules): Ingests newly disclosed agentic CVEs via Tavily and auto-synthesizes YAML rules using NVIDIA Nemotron.
  • Enterprise Reporting: Emits SARIF v2.1.0 for native GitHub Code Scanning integration, Rich terminal UI cards, and standalone HTML dashboards.

⚙️ How We Built It

We engineered AgentGuard as a high-throughput, 3-tier hybrid pipeline costing less than $0.003 per scan:

  1. Deterministic Rules Engine (Tier 1 - 0ms, $0.00): A Python-based AST and file walker traversing repositories while filtering .git, node_modules, and binary assets. It evaluates an 11-category YAML rule library grounded in real CVEs (Antigravity RCE, Cursor CVE-2024-51782, POSIX wildcards, Socket NPM lifecycle traps).
  2. Semantic Intent Layer (Tier 2 - NVIDIA Nemotron 3 Ultra 550B): Driven by NVIDIA's flagship nvidia/nemotron-3-ultra-550b-a55b via the OpenAI-compatible API. We optimized inference with a strict 100-token completion limit, disabled thinking overhead, and SHA-256 content deduplication caching.
  3. Live Threat Intelligence (Tier 3 - Tavily Search API): Extracts external URLs, endpoints, and shell commands, querying Tavily's threat intelligence with domain caching and a strict 3-query safety budget per scan.
  4. OS Runtime Interceptor: Implemented as a lightweight binary (guard) that evaluates process arguments and intercepts execution before child processes spawn.

🧪 Real-World Benchmarks & Challenges We Faced

  • The False-Positive Crisis: Early prototypes flagged legitimate .cursor/rules/*.mdc files because they contained instructions directed at an AI. We resolved this by engineering a Differential Intent Prompt for NVIDIA Nemotron that explicitly evaluates subversive authorization bypasses vs. benign architecture guidelines. On our voice-HHgoa testbed, false alarms dropped from critical warnings to zero.
  • Enterprise Scale: We benchmarked AgentGuard against a commercial 1,053-file full-stack repository (water-project). AgentGuard scanned all 1,053 files in 33.47 seconds with 0 false positives, proving enterprise readiness without slowing down developers.
  • Air-Gapped & Offline Execution: Enterprise teams require zero cloud dependency. We architected AgentGuard to run 100% offline via --skip-semantic on CPU, or by pointing the base URL to local Ollama / vLLM / NVIDIA NIM instances on local RTX GPUs.

🏆 Accomplishments That We're Proud Of

  • 11/11 Unit Tests Passing in 0.066s: Full coverage of the firewall, rule compilation, lockfile integrity, and kill-chain correlation.
  • Real OS Hard-Blocking: Achieving clean, non-bypassable process interception with exit code 126 and high-visibility Rich terminal alert panels.
  • GitHub Code Scanning Compatibility: Generated compliant SARIF v2.1.0 output directly ingestible into GitHub CodeQL security tabs.
  • Open-Source Delivery: Pushed clean, atomic Conventional Commits to our public repository with zero secret leaks.

🔮 What's Next for AgentGuard

  • IDE Plugins & Extensions: Direct integrations for Cursor, VS Code, and Windsurf to block prompt injections in real time during live agent chats.
  • MCP Server Interceptor: Native middleware proxy for the Model Context Protocol (MCP) to enforce permission boundaries on third-party agent tools.
  • Decentralized Threat Signatures: Community-contributed CVE rule registry continuously updated via our autonomous Tavily + Nemotron ingestion pipeline.

Built With

  • ai-agents
  • bash
  • claude-code
  • cursor
  • cybersecurity
  • devsecops
  • firewall
  • llm
  • nemotron
  • nvidia
  • prompt-injection
  • python
  • sarif
  • tavily
Share this project:

Updates

Submission history