Inspiration
Modern production incidents can escalate from a small anomaly into a major outage within minutes. SRE teams often have to correlate logs, metrics, chat messages, deployments, and customer impact while simultaneously communicating with stakeholders and documenting the incident.
We wanted to explore whether a group of specialized AI agents could assist with this entire incident lifecycle instead of treating incident response as a single chatbot task. That idea led us to build SRE-Brain, an AI-powered Site Reliability Engineering platform for automated incident detection, triage, communication, and post-mortem generation.
What it does
SRE-Brain processes simulated infrastructure telemetry such as latency, CPU usage, database connections, error rates, request rates, logs, and Slack-style incident context.
The system uses three specialized agents:
- TelemetryTriageAgent detects anomalies using sliding-window Z-score analysis, builds an incident timeline, and generates an AI-assisted root-cause hypothesis.
- CommsSyncAgent compresses noisy Slack context and generates stakeholder updates in different tones such as calm, technical, and crisis.
- PostMortemAgent combines incident evidence, GitHub context, financial impact, and AI-generated lessons learned into a structured Markdown post-mortem.
A graph-based router connects these agents so the incident can move through detection, investigation, communication, resolution, and documentation as one workflow.
The platform also includes a financial-impact calculator that estimates the cost of an outage in real time for each simulated scenario.
How we built it
We built the system in Python with a modular multi-agent architecture.
The agent workflow is implemented as a cyclic execution graph with shared incident state and Pydantic models. Statistical anomaly detection is handled using Z-score analysis with both predefined baselines and adaptive recent observations.
For AI reasoning, we integrated Google Gemini for root-cause hypotheses and incident-specific lessons learned. The system also supports rule-based fallbacks, so the core workflow can still operate without an API key.
We added GitHub and Slack integration adapters with mock/live modes, a Streamlit dashboard for the interface, and an evaluation suite covering detection precision/recall/F1, mean time to detect, financial accuracy, anomaly detection, and end-to-end graph execution.
The complete application can also be containerized and run with Docker.
Challenges we ran into
One of the biggest challenges was designing a reliable workflow around multiple AI agents without making the system dependent on free-form LLM output.
We separated deterministic components from AI reasoning: statistical anomaly detection, state management, routing, and financial calculations remain programmatic, while Gemini is used where language understanding and reasoning provide the most value.
Another challenge was handling noisy incident communication. Real incident channels can contain large amounts of panic, repetition, and irrelevant discussion, so we built a context-compaction layer to extract the signals most useful for diagnosis and communication.
We also had to make the system testable. Instead of relying only on subjective LLM output, we created simulated incident scenarios and quantitative benchmarks so that important parts of the pipeline could be measured.
Accomplishments that we're proud of
We are proud of turning incident response into an end-to-end multi-agent workflow rather than a simple AI assistant.
The platform can connect anomaly detection, timeline construction, AI-assisted RCA, communication drafting, financial impact estimation, and automated post-mortem generation in a single pipeline.
We also built an evaluation framework that measures technical performance such as alert detection precision/recall/F1, mean time to detect, financial calculation accuracy, and end-to-end execution correctness.
Most importantly, the architecture is modular: individual agents, tools, integrations, and evaluation components can be extended independently.
What we learned
We learned that useful AI systems for infrastructure should combine deterministic engineering with probabilistic AI reasoning.
LLMs are valuable for interpreting context, generating hypotheses, and producing human-readable reports, but they work better when surrounded by structured data, explicit state, deterministic tools, validation, and measurable evaluation.
We also learned that multi-agent systems are most useful when each agent has a focused responsibility and the overall workflow is explicitly orchestrated.
What's next for SRE-Brain
The next step is moving from simulated incidents toward deeper integration with real observability and incident-management systems.
We want to connect SRE-Brain with production telemetry sources, real Slack/incident channels, monitoring platforms, deployment systems, and additional runbooks. We also plan to strengthen root-cause validation, add more incident scenarios, improve evaluation coverage, and introduce stronger safeguards before allowing any automated remediation actions.
Our long-term goal is to build an AI SRE copilot that helps engineers detect, understand, communicate, and learn from incidents faster while keeping humans in control of high-impact actions.
Log in or sign up for Devpost to join the conversation.