Inspiration

AI coding assistants have made writing software dramatically faster than verifying it — and security is where that gap hurts most. Developers now merge code they didn't fully write, while the scanners meant to catch problems drown them in false positives. The result is the worst of both worlds: warnings everyone ignores, and confident-but-wrong misses that ship to production. Small teams and student developers are hit hardest — they rarely have a security reviewer at all.

We wanted to build the tool we wished existed: an agent that treats every security finding as a hypothesis, not a conclusion — and doesn't rest until the exploit either fires or it doesn't. One that proves vulnerabilities instead of guessing at them, fixes what it finds, and is honest about what it couldn't verify instead of quietly calling it safe.

What it does

Rouge is an autonomous red-team agent that finds vulnerabilities, proves them exploitable, ranks them, and patches them — end to end.

It runs a four-stage pipeline:

  1. Recon (The Scout) reads the entire target codebase and traces user input to dangerous sinks, producing candidate findings for SQL injection, command injection, XSS, path traversal, SSRF, and hardcoded credentials.
  2. Exploit (The Rouge Agent) writes a real test script for each finding and executes it against an isolated copy of the target. A vulnerability is CONFIRMED only when the script exits cleanly and prints a VULNERABLE: verdict with the exploit output as evidence. A tested-and-blocked finding becomes a FALSE POSITIVE. Anything that can't be verified — missing dependency, timeout, unsafe script — is marked UNVERIFIED, never silently counted as safe.
  3. Triage (The Triage Agent) ranks confirmed issues by exploitability, impact, and exposure, so the most dangerous problem gets fixed first.
  4. Patch (The Sentinel Agent) rewrites the vulnerable function, then re-runs the original exploit against the patched code. A fix is only accepted when the exploit is provably blocked (PATCH_VERIFIED:).

Every scan ends in one of three honest states: confirmed vulnerabilities with working exploits and verified patches, a VERT — Clean Bill of Health, or an INCOMPLETE verdict when findings exist that couldn't be tested. Everything streams live to a built-in monitoring dashboard, so you can watch each agent work in real time.

Setup takes one command (rouge setup — pick a model provider, paste a key, it validates and saves), and the model layer runs on any of 100+ providers, defaulting to Qwen.

How we built it

Rouge is a Python agent system orchestrated with LangGraph: four specialist agents (recon, exploit, triage, patch) as nodes over shared typed state, with a conditional edge routing to patch mode or report mode after triage.

  • A verification-first verdict model — the three-state result (confirmed / negative / inconclusive) is the spine of the system. Nothing downstream can promote an unverified finding to "safe," and reports carry a dedicated Unverified section explaining exactly why each test couldn't run.
  • A safety harness for LLM-generated code — exploit scripts are authored by a model that just read untrusted third-party source, so we don't trust them. Every script passes an AST safety gate (no network, subprocess, file writes, eval/exec), runs with secrets scrubbed from the environment, and executes in a disposable snapshot of the target — the user's repo is never mutated, and side effects like test databases land in a throwaway copy.
  • Provider-agnostic model layer — one LiteLLM client behind a clean interface, so the same agent graph runs on Qwen, OpenAI, Anthropic, Gemini, Groq, or a local Ollama model unchanged. A setup wizard with live credential validation writes to an owner-only config file.
  • A monitoring station — an aiohttp server with Server-Sent Events, a scan-launch API, and a live dashboard showing the agent log, the findings/confirmed/false-positive/unverified summary, and VERT/INCOMPLETE status.
  • Careful patch surgery — AST-based function replacement with line-proximity disambiguation for same-named functions, decorator preservation (so an @app.get route survives its own fix), and full-tree overlays for validation so nested packages and relative imports still resolve.

AI disclosure: LLMs are the reasoning engine of the project itself — every agent (recon analysis, exploit authoring, triage ranking, patch generation, remediation text) is an LLM call routed through LiteLLM, tested primarily with Qwen models. AI coding assistance was also used during development (implementation, debugging, and code review alongside handwritten code). All agent prompts, verdict logic, and safety controls were designed and verified by the team, including adversarial self-tests of the verification protocol.

Technologies: Python, LangGraph, LiteLLM (Qwen default), ast module, aiohttp/SSE, Click + Rich (CLI).

Challenges we ran into

The self-attestation trap. Our first exploit agent confirmed a vulnerability whenever the generated script exited with code 0. A script that just printed hello "confirmed" everything. The fix reshaped the whole system: a strict marker protocol (VULNERABLE: / NOT_VULNERABLE: / PATCH_VERIFIED:), a check that the script actually imports and calls the reported target, and a refusal to accept a bare zero exit code as proof. Then we applied the same rigour to patch validation, which had the identical flaw.

Fail-open was the real enemy. Worse than a wrong answer was a wrong silence. A missing fastapi import made an exploit script crash — and the scan reported a clean bill of health for a codebase full of SQL injection. We introduced the UNVERIFIED state and the INCOMPLETE verdict specifically so "we couldn't test this" can never masquerade as "this is safe," plus a dependency pre-flight that warns before the scan even starts.

Sandboxing code written by a model that just read hostile input. A scanned repository can contain prompt-injection payloads aimed at the exploit agent. We can't claim a true security boundary — and we're honest about that in the code — but we built meaningful guardrails: static AST checks, environment scrubbing so API keys never reach generated scripts, and filesystem isolation. During testing we even caught our own harness recursively snapshotting a temp directory and filling a disk, which taught us to exclude our own artifacts.

Ranking that actually ranked. Our triage step fed the model's ranking straight into sorted() as a sort key, which quietly discarded it. The shipped version converts the ordered output into a real index-position map, with a severity-based fallback for malformed responses.

Patching without breaking surrounding code. Replacing a function by string-matching def name( collides with same-named methods in other classes and matches inside comments. Moving to the ast module gave us exact spans, proximity-based disambiguation, and correct decorator handling.

Accomplishments that we're proud of

  • A confirmation loop that actually confirms. Exploit, then re-exploit-the-patch, with real stdout as evidence — it cleanly separated true positives from false positives on our test targets instead of guessing.
  • Verdicts we'd stake our name on. Three states, not two. The honest INCOMPLETE result is the design decision we're proudest of: we'd rather tell a developer "we couldn't verify this, look here" than hand them a false all-clear.
  • Safety we didn't hand-wave. An AST gate, secret scrubbing, and disposable execution for LLM-authored code — with the limitations documented in the code rather than hidden.
  • Adversarial self-testing. We proved the flaws empirically before fixing them: a no-op script confirming a finding, a malformed finding crashing the scan, a missing dependency producing a fake clean bill. Every fix has a regression test behind it.
  • Built for real users, not just judges. A one-command setup wizard with live validation, a live monitoring dashboard, structured reports and agent logs — the kind of thing that's meant to be run every day, including by developers with no security background.

What we learned

  • Never trust an agent's self-report. A model saying "this is vulnerable" or "the patch works" is a claim, not evidence. Evidence is a process that ran, output we can cite, and a verdict that could have come out the other way.
  • Fail-closed beats fail-silent. The hardest bugs weren't crashes — they were the places where an error path resolved to the safe-looking answer. Designing for "unknown is not safe" changed the architecture, not just the error handling.
  • Agents that run code need a harness, not a prayer. Prompting a model "don't use os.system" is not a control. Parsing its output before you execute it is.
  • Orchestration is where the quality lives. The individual prompts were the easy part. The value came from the contracts between agents — what counts as confirmed, what may be promoted, what must stay undecided.
  • Honesty is a feature. Telling a user "I don't know, here's why, here's what to do" turned out to be the most differentiating thing in a field full of tools that always sound certain.

What's next for Rouge

  • Real isolation. Move generated-code execution into containers or microVMs (seccomp, user namespaces) so the sandbox is a boundary rather than guardrails — the honest next step we already called out in the code.
  • Human-in-the-loop checkpoints. Turn the monitoring station into an approval gate: pause before a patch is applied, let a reviewer accept or reject, then resume. Autonomy with a seatbelt.
  • CI/CD integration. A GitHub Action that runs Rouge on every pull request, posts confirmed findings with exploit evidence as a review comment, and proposes the verified patch as a suggested change — catching vulnerabilities before they merge, where they're cheapest to fix.
  • Broader coverage. More languages (JavaScript/TypeScript web apps next), more vulnerability classes (auth bypass, insecure deserialization, secret leakage), and framework-aware analysis.
  • Memory and learning. Let the agent accumulate which exploit patterns actually succeed against a codebase over time, so recon gets sharper with every scan.
  • Hosted scanning. A persistent service with a scan queue and stored reports, so a whole team or classroom can point Rouge at repos without running anything locally.

Rouge is our answer to a simple question: if AI is going to write our software, can it at least prove it's safe? We think the answer is yes — one exploit at a time.

Built With

Share this project:

Updates

Submission history