About the project Production incidents are rarely caused by a lack of alerts—they are prolonged by a lack of context. During an outage, engineers jump between dashboards, logs, recent deployments, runbooks, tickets, and tribal knowledge to answer basic questions: Has this happened before? What fixed it? Who owns the affected service?
That problem inspired Aegis Swarm: a memory-first, multi-agent incident commander. It turns a production alert into a coordinated investigation, using the lessons of previous incidents to help responders move from detection to safe resolution faster.
What it does Aegis Swarm ingests an incident or alert, retrieves semantically similar historical incidents, and coordinates specialist agents across the investigation. Each agent focuses on a specific responsibility—such as log analysis, deployment correlation, dependency analysis, runbook retrieval, or incident summarization—while sharing findings through a common incident context.
Instead of giving engineers another generic chat interface, Aegis aims to provide: Relevant historical incidents and their outcomes Evidence-backed hypotheses about likely causes A live, auditable incident timeline Suggested next steps and runbooks Human-approved remediation guidance
How I built it I designed Aegis Swarm around a durable shared-memory layer powered by CockroachDB and deployed its agentic workflows on AWS.
The system stores incident records, service context, runbooks, investigation artifacts, agent findings, and resolution outcomes in a structured operational memory. When a new alert arrives, Aegis retrieves related incidents using semantic search and combines them with current telemetry and service metadata.
A coordinator agent then delegates tasks to specialized agents. Their findings are written back to the shared incident context, allowing later agents—and the human responder—to work from the same evolving picture of the incident.
I used CockroachDB’s distributed database capabilities as the foundation for resilient incident data and coordination, while integrating agent tooling and cloud services on AWS to support the workflow.
Challenges I faced The hardest challenge was making a multi-agent system useful rather than noisy. Multiple agents can easily generate overlapping conclusions, conflicting recommendations, or unsupported guesses. We addressed this by emphasizing shared context, evidence tracking, scoped agent roles, and a human-in-the-loop approach for impactful actions.
Another challenge was modeling institutional knowledge. Incident reports, runbooks, and postmortems are often inconsistent or incomplete. I had to think carefully about how to preserve source context, connect related records, and surface historical lessons without presenting them as guaranteed answers.
Finally, incident response is a high-trust environment. Aegis must be transparent about what it knows, what it infers, and where uncertainty remains. The goal is not to replace the incident commander—it is to give them the context and coordination capacity of a much larger response team.
What I learned I learned that reliable agent systems depend as much on memory design and orchestration as on model capability. Retrieval quality, clear task boundaries, durable shared state, and traceable evidence are essential when agents operate in operationally sensitive workflows.
I also learned that the most valuable AI assistance is not an answer in isolation. It is the ability to connect a current incident to the organization’s prior decisions, failures, and recoveries—then help humans act on that context with confidence.
Built With
- ai
- amazon-web-services
- automation
- cloud
- cockroachdb
- computing
- cursor
- cybersecurity
- database
- developer
- devops
- distributed
- docker
- engineering
- llm
- multi-agent
- observability
- python
- rag
- reliability
- search
- semantic
- systems
- vector
Log in or sign up for Devpost to join the conversation.