Inspiration

Cybersecurity teams are overwhelmed by the volume of threat intelligence, vulnerability information, security alerts, and attack techniques they need to analyze. Traditional workflows often require analysts to manually search multiple sources, correlate evidence, assess risk, and determine appropriate remediation.

Our previous work on an AI-powered Cyber Threat Intelligence platform showed us that Retrieval-Augmented Generation can make security investigation more contextual and useful. The platform combines hybrid retrieval using FAISS and BM25 with an LLM to support threat intelligence querying, CVE analysis, threat classification, risk scoring, and security reporting.

CyberShield AI builds on that foundation with a more agentic approach: instead of treating an LLM as only a question-answering interface, we are designing a security agent that can reason over evidence, investigate threats, determine risk, and assist with remediation workflows.

What We Built

CyberShield AI is an AI-powered cybersecurity agent designed around the workflow:

Detect → Investigate → Reason → Assess → Remediate → Validate

The system combines:

  • AI-powered security reasoning
  • Retrieval-Augmented Generation
  • Hybrid semantic and keyword retrieval
  • Threat intelligence analysis
  • Vulnerability intelligence
  • Risk assessment
  • Security recommendations
  • Automated security reporting

The underlying CTI foundation uses FAISS + BM25 for hybrid retrieval and Ollama with Qwen2.5-Coder 7B for local LLM-based reasoning.

The architecture is designed to remain modular so that additional security-analysis tools, datasets, and agent capabilities can be incorporated as the project evolves.

How It Works

1. Security Input

The system receives security information such as threat intelligence, vulnerability information, suspicious activity, or security-related code.

2. Context Retrieval

Instead of relying only on the LLM's internal knowledge, relevant security context is retrieved using a hybrid retrieval pipeline.

FAISS provides semantic retrieval while BM25 provides keyword-based retrieval.

3. AI Security Reasoning

The retrieved evidence is provided to the LLM to generate a contextual security analysis.

The reasoning process can identify:

  • Relevant threats
  • Vulnerability context
  • Attack techniques
  • Potential impact
  • Risk factors
  • Recommended security actions

4. Risk Assessment

The system evaluates the available evidence and produces a structured risk assessment to help prioritize security findings.

5. Remediation Assistance

Based on the analysis, the system can generate actionable security recommendations and remediation guidance.

6. Security Reporting

The investigation results can be converted into structured intelligence outputs and security reports for analysts.

Agentic Approach

The key idea behind CyberShield AI is moving beyond a simple chatbot.

A conventional LLM workflow looks like:

User → Prompt → LLM → Response

CyberShield AI is designed around a security workflow:

Security Event → Retrieve Evidence → Analyze → Reason → Assess Risk → Recommend Action → Validate

This allows the AI system to operate as a security reasoning layer rather than simply generating conversational responses.

Technology Stack

AI / Agent Layer

  • Python
  • LangChain
  • Ollama
  • Qwen2.5-Coder 7B
  • Retrieval-Augmented Generation

Retrieval

  • FAISS
  • BM25

Backend

  • FastAPI
  • SQLAlchemy
  • PostgreSQL
  • JWT Authentication

Frontend

  • Next.js
  • TypeScript
  • Axios

Deployment

  • Vercel
  • Railway
  • PostgreSQL
  • GitHub

Challenges

One of the main challenges was making LLM-based cybersecurity analysis grounded in reliable context rather than allowing the model to generate unsupported conclusions.

This motivated the use of hybrid retrieval instead of relying solely on an LLM. Combining semantic retrieval through FAISS with keyword retrieval through BM25 allows the system to retrieve both conceptually relevant and explicitly matching security information.

Another challenge was designing the system so that security analysis remains modular. Threat intelligence, vulnerability intelligence, retrieval, AI reasoning, risk scoring, and reporting need to work as separate components while still forming a single investigation workflow.

A further challenge is validation. An AI-generated recommendation should not automatically be treated as a successful security fix. Future iterations of CyberShield AI are designed to incorporate stronger automated testing and validation so that remediation can be evaluated rather than simply generated.

What We Learned

Building the CTI foundation taught us that integrating an LLM into cybersecurity is not simply a matter of sending security questions to a model.

The quality of the result depends heavily on:

  • Relevant context
  • Retrieval quality
  • Evidence grounding
  • Structured reasoning
  • Security-specific workflows
  • Validation of generated recommendations

We also learned that agentic systems are most useful when they are connected to real tools and evidence sources rather than operating as isolated conversational models.

Future Direction

Our next development stage is to expand CyberShield AI toward a more autonomous security workflow where specialized agents and security-analysis tools can collaborate.

The planned workflow is:

Detect → Retrieve → Investigate → Reason → Assess → Remediate → Test → Validate

This architecture can be extended with static analysis, dynamic analysis, fuzzing, additional threat-intelligence sources, and security validation tools.

The goal is not to replace security analysts, but to reduce the time required to move from a raw security finding to an evidence-backed and actionable security decision.

Links

GitHub: https://github.com/sayeedur007-design/Cyber-Threat-Intel

Live Platform: https://cyber-threat-intel-qu0pvd81j-sayeedur007-designs-projects.vercel.app/login

Built With

Share this project:

Updates

Submission history