Inspiration

Software teams usually discover security vulnerabilities too late—after deployment, during a penetration test, or when a real attacker finds them first.

Traditional code scanners are useful, but they mostly look for predefined patterns. AI code reviews often provide broad recommendations without proving whether an issue is actually exploitable.

We wanted to build something more proactive: a swarm of autonomous AI agents that could think and operate like attackers, challenge a codebase from multiple angles, verify each other’s findings, and help developers hack their own product before someone else does.

What it does

Agent Swarm is an autonomous cybersecurity testing platform that launches a coordinated team of specialized AI agents against a software project.

The system first maps the repository, identifies the technology stack, and discovers high-risk attack surfaces. It then creates an attack plan and routes each task to the most relevant specialist agents, including:

  • Authentication and authorization breakers
  • API abuse specialists
  • Secrets and configuration auditors
  • Repository reconnaissance agents
  • Vulnerability verification agents

Instead of relying on a single AI response, agents independently investigate different attack paths. Their findings are then verified, deduplicated, ranked by severity, and combined into a clear security report.

Agent Swarm can analyze source code, configurations, dependencies, APIs, and application architecture while operating under configurable safety boundaries.

How we built it

We built Agent Swarm as a Python CLI application with LangGraph coordinating the multi-agent workflow.

The pipeline is divided into several stages:

  1. Repository intelligence inventories files, detects the stack, and identifies potential attack surfaces.
  2. An agentic retrieval system selects the most relevant cybersecurity techniques from a library of over 800 security skills.
  3. A reconnaissance agent summarizes the application and its exposed surfaces.
  4. An attack planner creates targeted security objectives.
  5. Specialized agents investigate authentication, APIs, secrets, configurations, and other risk areas.
  6. A verifier challenges each proposed vulnerability and removes weak or unsupported findings.
  7. The final ranking system deduplicates issues and prioritizes them by severity, confidence, and potential impact.

We also added provider configuration, local caching, environment validation, CLI onboarding, configurable risk levels, and strict target restrictions so developers can control how aggressively the swarm operates.

Challenges we ran into

One of the hardest problems was preventing the agents from producing plausible-sounding but unverified vulnerabilities. Cybersecurity findings need strong evidence, not speculation, so we designed a dedicated verification stage that checks each claim against the repository and rejects unsupported results.

Another challenge was efficiently routing more than 800 cybersecurity skills. Sending every skill to every agent would create excessive context and reduce performance. We built an agentic retrieval layer that dynamically selects only the techniques relevant to the current stack, attack surface, and objective.

Coordinating multiple agents also introduced duplicate findings and conflicting conclusions. We created structured schemas, deterministic fingerprints, deduplication logic, and confidence-based ranking to consolidate overlapping results.

Finally, we had to balance autonomy with safety. The system needed to behave creatively like an attacker while remaining restricted to explicitly approved local targets and configurable risk levels.

Accomplishments that we're proud of

We are proud that Agent Swarm goes beyond a basic vulnerability scanner or single-agent code review.

The system can autonomously move from an unfamiliar repository to a structured attack plan, select relevant security knowledge, coordinate specialized agents, verify evidence, and produce prioritized findings.

We successfully integrated a cybersecurity knowledge base containing over 800 skills without placing the entire library into every agent’s context.

We also built a modular architecture where reconnaissance, planning, specialization, verification, deduplication, ranking, and model providers can evolve independently.

Most importantly, Agent Swarm demonstrates how multi-agent systems can simulate an actual security team rather than simply generate another static analysis report.

What we learned

We learned that more agents do not automatically produce better results. Each agent needs a narrow role, clear tools, structured outputs, and a reason to exist within the workflow.

We also learned that retrieval quality is more important than context quantity. Giving an agent a small set of highly relevant security techniques produced better results than providing a massive unfiltered knowledge base.

Verification turned out to be just as important as discovery. An autonomous security system must be able to question its own conclusions and distinguish real vulnerabilities from theoretical concerns.

Finally, we learned that effective agent systems require both deterministic software and probabilistic reasoning. The agents provide creativity and adaptability, while schemas, validation, deduplication, and safety controls provide reliability.

What's next for Agent Swarm

Next, we plan to expand Agent Swarm from static repository analysis into controlled runtime security testing.

Future improvements include:

  • Browser and API-based dynamic testing
  • Sandboxed exploit validation
  • GitHub Actions integration
  • Pull-request security reviews
  • Additional specialist agents
  • Attack-path visualization
  • Support for custom organization-specific security skills
  • Continuous learning from verified findings and developer feedback

Our long-term goal is to make Agent Swarm an always-available autonomous security team that continuously challenges software throughout the development lifecycle—before real attackers get the opportunity.

Built With

Share this project:

Updates

Submission history