Inspiration

Security auditing for modern software is fundamentally broken. It relies heavily on static analysis tools that throw thousands of false positives, or expensive human pentesters who can't scale to match the speed of CI/CD pipelines. We were inspired by the concept of "Agentic Societies"—what if, instead of a single monolithic AI trying to find bugs, we built an entire specialized civilization of agents?

We wanted to mimic a real-world cybersecurity firm: a scout finding targets, researchers mapping attack surfaces, elite hackers writing exploits, skeptical auditors verifying them, and engineers writing the patches. NEXUS was born out of the belief that specialized, communicating agents governed by a consensus model can achieve near-human precision in vulnerability research.

What it does

NEXUS is a self-evolving, multi-agent security civilization powered by Alibaba Cloud. When you point NEXUS at a GitHub repository, an orchestrated pipeline of 10 distinct Qwen-powered agents springs into action:

  1. Scout: Crawls the repository to map its file structure and architecture.
  2. Recon: Analyzes the attack surface, isolating high-risk entry points like auth middleware and SQL queries.
  3. Hunter: Performs deep codebase analysis using Qwen-Max's massive context window to detect vulnerabilities.
  4. Exploit: Generates safe, proof-of-concept (PoC) code to prove the vulnerability is real.
  5. Verify: Acts as a skeptic, cross-validating the PoC to eliminate false positives.
  6. Governance (3 Agents): A council of three distinct personas (CVSS Scorer, Impact Assessor, Exploitability Judge) that debate and vote to reach a consensus on the vulnerability's severity.
  7. Patch: Generates an AST-aware security fix with regression tests.
  8. Review: Ensures the patch is clean, correct, and doesn't introduce breaking changes.
  9. Report: Generates a CVE-ready security advisory.

Finally, NEXUS publishes the scan reports and exploit artifacts directly to Alibaba Cloud OSS for immutable storage and sharing.

How we built it

NEXUS is built entirely on Alibaba Cloud infrastructure, leveraging DashScope API to orchestrate the intelligence. We used Qwen-Max for high-reasoning tasks (like hunting vulnerabilities and generating exploits) and Qwen-Plus for broader context tasks (like scouting and reporting) to optimize cost and latency.

The architecture consists of:

  • Backend: A FastAPI Python server utilizing ARQ and Redis for background task queuing. The orchestration layer connects the 10 agents and isolates them so a failure in one doesn't crash the pipeline.
  • Memory System: A 3-tier memory engine:
    • L1 Working Memory (Redis): Real-time agent communication and context.
    • L2 Episodic Memory (PostgreSQL): Permanent storage of scan results and governance votes.
    • L3 Semantic Memory (pgvector): Stores vulnerability pattern embeddings so NEXUS learns over time.
  • Frontend: A Next.js enterprise-grade dashboard featuring a real-time WebSocket terminal that streams agent thoughts, live metrics, and a dynamic vulnerability Kanban board.
  • Cloud Infrastructure: Deeply integrated with Alibaba Cloud OSS for artifact storage via the oss2 SDK, designed to be deployed on Alibaba Cloud ECS.

Challenges we ran into

1. The "Yes Man" AI Problem: Initially, our Verify agent would just agree with whatever the Hunter agent found. We had to heavily engineer the system prompts to create a genuinely skeptical persona whose only job was to try and disprove the Hunter.

2. Managing Chaos in a 10-Agent Pipeline: Ensuring structured JSON output from LLMs across a complex pipeline was extremely fragile. If one agent hallucinated a malformed JSON string, the whole chain broke. We solved this by implementing strict Pydantic model validation boundaries and per-agent error isolation.

3. Real-Time State Sync: Getting 10 background AI agents to smoothly report their thought processes and status to a React frontend in real-time required building a custom Redis PubSub event bus connected to FastAPI WebSockets.

Accomplishments that we're proud of

  • The Governance Council: We are incredibly proud of the 3-agent voting mechanism. Watching three separate LLM personas debate a vulnerability and mathematically average their CVSS scores to reach a consensus feels like a glimpse into the future of autonomous organizations.
  • Zero False Positives by Design: By forcing the system to generate a PoC (Exploit agent) and independently verify it (Verify agent), we shifted from a model of "guessing" bugs to cryptographically proving them.
  • The UI Experience: We built a dashboard that doesn't just look like a standard web app—it feels like an enterprise security command center, making the invisible work of AI agents visceral and transparent.

What we learned

  • Specialization > Generalization: Splitting a complex task (security auditing) into micro-roles drastically improves the LLM's performance. A single prompt asking an LLM to "find a bug and fix it" fails constantly; a pipeline of 10 specialized prompts works beautifully.
  • Alibaba Cloud Ecosystem: We learned how seamless it is to integrate DashScope's OpenAI-compatible endpoints with existing orchestration frameworks, and how robust the oss2 SDK is for managing artifacts on the fly.

What's next for NEXUS

NEXUS v1.0 proves that an agent society can audit code, but v2.0 will actively defend it. Our next steps include:

  1. Auto-PR Generation: Automatically branching the repo, applying the Patch agent's fix, and opening a Draft PR on GitHub.
  2. Semantic Forgetting and Evolution: Enhancing the L3 memory so that when NEXUS successfully exploits a bug, it generates an embedding of that code pattern and actively hunts for similar zero-days in other repositories.
  3. Live Sandbox Execution: Integrating Alibaba Cloud ACK (Kubernetes) to safely spin up the target repository in a live sandbox, allowing the Exploit agent to actually run its PoC scripts against the live codebase.

Built With

  • alibaba
  • cloud
  • dashscope
Share this project:

Updates

Submission history