Inspiration
Cloud incidents rarely happen at convenient times. When a production service starts returning 502 errors, an SRE may need to inspect CPU and memory usage, search logs, check dependent services, compare the incident with previous failures, decide whether a restart is safe, and finally execute the remediation. Each step takes time and during an outage, every minute matters.
We wanted to explore a different approach: What if an AI agent could act like an SRE, investigate an incident autonomously, learn from previous incidents, and safely remediate the problem—while keeping humans in control of high-risk actions?
That idea became HealOps, an autonomous AI SRE designed around a simple principle:
Detect → Diagnose → Decide → Heal → Learn.
Instead of building another monitoring dashboard that only tells an engineer something is wrong, we built a system that can investigate what happened, reason about the likely cause, propose an action, enforce safety policies, and execute remediation when permitted.
What it does
HealOps is an autonomous AI Site Reliability Engineer and incident auto-remediation platform.
It provides an SRE Mission Control dashboard where operators can monitor simulated infrastructure, trigger controlled failures, watch the AI agent investigate the incident in real time, approve remediation when required, and review the resulting audit trail.
The core incident workflow is:
- Detect — HealOps receives telemetry and identifies an anomalous service.
- Investigate — The Strands-based AI agent uses diagnostic tools to inspect system resources, services, logs, HTTP health, databases, and Redis.
- Reason — The agent correlates observations and compares the failure signature with historical post-mortems/runbooks.
- Decide — The proposed remediation is evaluated through safety guardrails and intervention policies.
- Human approval — High-impact actions can be intercepted and presented to an operator for approval.
- Remediate — Once approved, HealOps can execute actions such as restarting a service or flushing a cache namespace.
- Verify — The system observes the infrastructure after remediation to determine whether the incident has been resolved.
- Audit & learn — Actions and decisions are recorded, while incident outcomes contribute to the system's historical runbook memory.
The platform also includes a Chaos Engineering Lab, allowing us to intentionally inject a simulated 502 outage and demonstrate the entire autonomous remediation loop end-to-end.
How we built it
HealOps was designed as a layered architecture combining an agentic AI system, infrastructure diagnostic tools, policy-based safety controls, and a real-time web interface.
Tech Stack
AI & Agent Layer
- AWS Strands Agents SDK
- Multi-model factory supporting Bedrock, Gemini, Claude, and Ollama
- Autonomous tool-calling and reasoning workflow
- Historical post-mortem/runbook memory
Backend & APIs
- Python 3.12+
- FastAPI
- REST APIs
- WebSockets for real-time event streaming
Infrastructure & Data
- Redis for agent sessions and event streams
strands-redis-session-manager- JSONL-based audit logging
- Docker/service registry integration
- SQLAlchemy for database health inspection
Observability & Diagnostic Tools
psutilfor system telemetryhttpxfor HTTP service probing- Service and container inspection
- Log inspection
- Database health checks
- Redis memory inspection
Safety & Governance
- Cedar-style policy engine
- Strands
InterventionHandler - Risk-based remediation policy matrix
- Human-in-the-Loop approval workflow
- SOC-2-oriented audit logging
Frontend
- HTML
- CSS
- JavaScript
- WebSocket-based live updates
- Dark-mode SRE Mission Control interface
The architecture separates the system into four major layers:
Layer 1 — Agent Brain: The Strands autonomous SRE agent handles investigation, reasoning, model selection, session state, and historical incident memory.
Layer 2 — Tool Suite: The agent interacts with controlled diagnostic and remediation tools rather than directly manipulating infrastructure.
Layer 3 — Safety & Governance: Potentially dangerous actions pass through policy enforcement and human approval before execution.
Layer 4 — Real-Time Gateway: FastAPI exposes the APIs and WebSocket stream that connect the autonomous backend with the SRE dashboard.
This separation allowed us to make the agent powerful without giving it unrestricted control over the infrastructure.
The complete system architecture is documented in the project repository.
Challenges we ran into
1. Making an AI agent useful beyond conversation
The biggest challenge was moving from a chatbot-style AI to an agent capable of performing an actual SRE workflow.
The agent needed access to structured tools for telemetry, logs, services, databases, HTTP endpoints, and remediation. We had to design these tools so the model could investigate systematically rather than simply generate a textual answer.
2. Balancing autonomy with safety
Giving an AI permission to restart services creates an important question:
How much autonomy is too much?
We designed HealOps so that the agent does not automatically execute every action it proposes. Remediation actions can pass through policy evaluation and a Human-in-the-Loop approval layer.
This became one of the central architectural decisions of the project: autonomy should increase operational speed without removing operational control.
3. Creating a convincing real-time experience
Incident response is inherently dynamic. A static dashboard would not communicate what an autonomous agent is actually doing.
We therefore built a WebSocket event stream that surfaces events such as incident detection, triage initiation, agent thoughts/actions, approval requests, and incident resolution directly in the dashboard.
4. Giving the agent memory
Diagnosing an incident from scratch every time is inefficient. We wanted HealOps to benefit from previous incidents.
We introduced a historical post-mortem/runbook memory layer so that the agent can compare current failure signatures with previously resolved incidents and use that context during investigation.
5. Demonstrating the system without risking real infrastructure
For a hackathon prototype, we needed a reliable way to demonstrate failure and recovery without causing an actual production outage.
The Chaos Engineering Lab solves this by providing a controlled 502 failure injection workflow. This lets us reproduce the incident, observe the agent's investigation, trigger the safety workflow, and demonstrate remediation safely.
Accomplishments that we're proud of
We are especially proud that HealOps is more than an AI-powered dashboard—it demonstrates an end-to-end autonomous incident response loop.
A complete autonomous SRE workflow
We implemented the complete flow from:
Chaos injection → incident detection → AI investigation → historical incident matching → policy interception → human approval → remediation → resolution → audit trail → post-mortem memory.
Agentic AI with real tools
Instead of using an LLM only to explain an incident, HealOps gives the agent an actual diagnostic toolbox and lets it decide which tools to use during investigation.
Human-governed autonomy
The project demonstrates that autonomous AI and human oversight do not have to be opposites. HealOps allows the AI to investigate independently while keeping operators in the loop for sensitive remediation decisions.
Real-time agent observability
The dashboard exposes the agent's operational workflow as it happens through WebSockets, making the otherwise invisible reasoning and execution process observable to an SRE.
Built-in chaos engineering
The integrated Chaos Engineering Lab gives us a repeatable way to test the system's incident-response capabilities rather than relying only on static demonstrations.
Auditability
Every important intervention can be captured through the audit layer, creating a trace of what the system observed, decided, and executed.
A production-inspired architecture
Although HealOps is a hackathon project, we intentionally designed the architecture around real SRE concerns: observability, controlled remediation, policy enforcement, human approval, persistent sessions, incident memory, and auditability.
What we learned
The biggest lesson was that building an autonomous agent is not just about choosing a powerful model.
The surrounding system matters just as much.
We learned that an effective infrastructure agent needs:
- Well-defined tools rather than unrestricted infrastructure access
- Structured telemetry and observable system state
- Persistent context and historical incident memory
- Explicit safety policies
- Human escalation paths
- Reliable event streaming
- Strong auditability
- Controlled environments for testing autonomous actions
We also learned that autonomy needs boundaries. In infrastructure operations, the goal isn't to give an AI unlimited authority. The goal is to give it enough capability to handle repetitive operational work while ensuring that risky decisions remain governed.
Building HealOps also changed how we think about AI agents: the most valuable agents are not necessarily the ones that generate the most impressive responses—they are the ones that can reliably complete meaningful workflows in the real world.
What's next for HealOps
The current version focuses on demonstrating the autonomous incident-response loop in a controlled environment. The next step is to evolve HealOps toward a production-ready autonomous SRE platform.
Production infrastructure integrations
We want to connect HealOps with real cloud and container environments such as Kubernetes, AWS infrastructure, and production observability platforms.
More remediation strategies
Beyond service restarts and cache flushing, future versions could support controlled actions such as rollback, scaling, configuration recovery, dependency isolation, and traffic shifting.
Smarter incident correlation
HealOps could correlate metrics, logs, traces, deployments, infrastructure changes, and historical incidents to improve root-cause analysis.
Adaptive risk policies
The policy layer could evolve toward more granular risk scoring, allowing HealOps to automatically execute low-risk actions while escalating increasingly dangerous operations to human operators.
Continuous learning from incidents
Resolved incidents could automatically become structured post-mortems and reusable runbooks, allowing the system's operational knowledge to grow over time.
Multi-agent SRE workflows
Future versions could introduce specialized agents for observability, diagnosis, security, remediation, and post-mortem generation, coordinated by an SRE orchestrator.
Our long-term vision is for HealOps to become a trusted autonomous operations layer—one that doesn't just tell engineers when infrastructure is broken, but actively helps restore it while remaining transparent, auditable, and human-governed.
Built With
- 3.12+
- amazon-web-services
- claude
- docker
- fastapi
- gemini
- javascript
- ollama
- python
- redis
- rest-api
- sqlalchemy
- strands-adk
Log in or sign up for Devpost to join the conversation.