Inspiration

SRE teams waste countless hours manually diagnosing repetitive cloud incidents and executing complex runbooks under high-pressure conditions. We wanted to build a system that brings intelligent, agentic automation to cloud operations safely.

What it does

ReplayOps is an intelligent, agentic runbook automation platform that detects cloud incidents, proposes root-cause analysis, and safely executes automated mitigation runbooks with human-in-the-loop controls.

How we built it

  • Backend: Built with high-performance FastAPI, handling asynchronous event processing, secure API routing, and AI agent orchestration.
  • Frontend: Crafted with a responsive React dashboard that displays real-time telemetry, incident logs, and interactive runbook approval prompts.
  • Research Grounding: Architected using automated SRE frameworks inspired by industrial research on human-in-the-loop agentic runbook improvements.

Challenges we ran into

Balancing fully autonomous AI execution with strict safety guardrails was tricky. We solved this by designing an explicit human-in-the-loop gating mechanism that requires engineer authorization before critical scripts execute.

Accomplishments that we're proud of

Successfully integrating an asynchronous Python agent pipeline with a live-updating React frontend, achieving verified response metrics that slash incident MTTR significantly.

What we learned

We gained deep production-level experience in designing reliable agentic workflows, managing stateful asynchronous operations in FastAPI, and building clean operator dashboards in React.

What's next for ReplayOps

Expanding integrations with more cloud monitoring tools and adding advanced multi-agent predictive incident forecasting.

Built With

Share this project:

Updates

Submission history