AlertPilot — Autonomous SOC Triage Agent Inspiration Security Operations Centers are drowning. The average enterprise SOC receives over 10,000 alerts per day — and more than 45% go uninvestigated due to analyst fatigue and bandwidth constraints. The result: real threats buried under noise, and burnt-out analysts making poor decisions at 2 AM. The root problem isn't intelligence. Modern threat intel APIs know exactly how dangerous an IP is. The problem is the gap between knowledge and action — someone still has to read every alert, look up every IP, decide the severity, and choose a response. That loop doesn't scale. AlertPilot closes that gap. It's not a dashboard. It's not a recommender. It's an autonomous agent that receives alerts, investigates them, decides what to do, and does it — while maintaining a full audit trail so human analysts stay in control.

What it does

AlertPilot is a hybrid-intelligence SOC agent built on Qwen that operates a continuous autonomous triage loop: Ingests security alerts — from its simulated live feed or via manual submission Enriches each alert in real time using AbuseIPDB and VirusTotal threat intelligence Reasons for the enriched context using Qwen to assign severity, generate analysis, identify false positive indicators, and produce remediation steps Acts autonomously based on its decision: CRITICAL → Auto-block IP + immediate escalation to analyst queue HIGH → Flag IP for monitoring + queue for human review MEDIUM → Add to watchlist + monitor for escalation LOW / INFO → Log to incident register Logs every decision to a tamper-evident audit trail with full reasoning Routes uncertain cases to a Human-in-the-Loop (HITL) queue where analysts can Approve, Reject, or Monitor The result: analysts spend their time reviewing agent decisions and handling escalations — not triaging noise.

How we built it

Backend: Python + FastAPI serving a fully async agent loop LLM: Qwen (via Alibaba Cloud DashScope) — structured JSON responses for severity, confidence, reasoning, false-positive indicators, and remediation steps Threat Intelligence: AbuseIPDB — IP reputation, abuse confidence score, Tor exit node detection, ISP identification VirusTotal — multi-engine malware detection, IP reputation scoring Database: SQLite with three tables: incidents — full alert + enrichment + analysis record per event audit\_log — every agent action timestamped with actor and reasoning hitl\_queue — pending human review items with analyst decision tracking Frontend: React single-page dashboard with five panels: Dashboard — live stats (total incidents, by severity, blocked IPs, pending review) Live Feed — real-time incident stream with severity, action taken, and expandable reasoning HITL Queue — analyst interface with Approve / Reject / Monitor controls Audit Trail — full chronological agent decision log Manual Triage — direct alert submission with immediate analysis and action output Agent Loop: Autonomous async loop (configurable interval, default 30s) that processes alerts end-to-end without human input. Fully restartable. Each cycle: generate → enrich → analyze → act → log.

Challenges we ran into

  1. Trust vs. Autonomy tradeoff The hardest design question was: what should the agent do on its own, and when should it stop and ask? We resolved this with a confidence threshold — HIGH severity incidents only trigger auto-block if Qwen's confidence exceeds 75%. Below that, the IP is flagged and escalated. CRITICAL always gets both automated action and human notification — because even correct autonomous decisions on critical threats need human awareness.
  2. Windows DNS resolution in async background tasks The FastAPI background agent loop had a different network context than the request-handling path on Windows, causing DNS resolution failures for AbuseIPDB and VirusTotal calls. Fixed by switching to absolute dotenv path resolution and per-call async HTTP clients with explicit timeouts.
  3. Structured LLM output reliability Qwen needs to return valid JSON every time — severity, confidence, reasoning, steps — and the pipeline breaks if it doesn't. We handled this with a strict system prompt, low temperature (0.2), JSON fence stripping, and a heuristic mock fallback that keeps the agent running during API unavailability.

Accomplishments that we're proud of

A fully autonomous agent loop that processes, classifies, and acts on security alerts without human input Real threat intelligence integration with live enrichment data driving Qwen's reasoning A working HITL queue with analyst override controls — the critical missing layer in most autonomous SOC tools A complete audit trail that explains every decision the agent made and why The dashboard going from zero to 36 incidents processed in under 30 minutes of autonomous operation during development

What we learned

Building an autonomous agent that people would actually trust in a security context is fundamentally different from building one that just works. Every design decision came back to the same question: what happens when it's wrong? The audit trail isn't a nice-to-have feature — it's the feature that makes the rest of the agent deployable. Without it, a SOC team would never hand autonomous action authority to an agent. With it, every block, every flag, every escalation is reviewable, reversible, and explainable. Qwen's reasoning quality is particularly strong for this use case. Given enriched threat intel data, it consistently identifies nuanced context — distinguishing Tor exit nodes used for legitimate privacy from the same infrastructure used for attacks, flagging false-positive indicators where benign services share ASNs with malicious actors — in a way that raw score thresholds can't match.

What's next for AlertPilot

Webhook ingestion — connect directly to SIEM outputs (Splunk, Elastic, Microsoft Sentinel) so AlertPilot receives real production alerts Action integrations — actual firewall API calls (pfSense, AWS Security Groups), Jira/ServiceNow ticket creation, Slack escalation notifications Multi-agent architecture — specialist sub-agents for different alert types (network, endpoint, cloud) coordinated by AlertPilot as orchestrator Analyst feedback loop — HITL decisions feed back into Qwen's prompting context, improving classification accuracy over time Multi-tenancy — organization-level incident isolation for MSSP deployment

Track Track 4 — Autopilot Agent: End-to-end autonomous business workflow automation. AlertPilot automates the full SOC alert triage workflow — from alert ingestion through threat investigation, decision-making, action execution, and audit logging — with human oversight through the HITL queue.

Built With

  • abuseipdb
  • alibaba-cloud
  • fastapi
  • httpx
  • openai-sdk
  • python
  • qwen
  • react
  • sqlite
  • uvicorn
  • virustotal-api
Share this project:

Updates

Submission history