Inspiration
SRE teams waste countless hours manually diagnosing repetitive cloud incidents and executing complex runbooks under high-pressure conditions. We wanted to build a system that brings intelligent, agentic automation to cloud operations safely.
What it does
ReplayOps is an intelligent, agentic runbook automation platform that detects cloud incidents, proposes root-cause analysis, and safely executes automated mitigation runbooks with human-in-the-loop controls.
How we built it
- Backend: Built with high-performance FastAPI, handling asynchronous event processing, secure API routing, and AI agent orchestration.
- Frontend: Crafted with a responsive React dashboard that displays real-time telemetry, incident logs, and interactive runbook approval prompts.
- Research Grounding: Architected using automated SRE frameworks inspired by industrial research on human-in-the-loop agentic runbook improvements.
Challenges we ran into
Balancing fully autonomous AI execution with strict safety guardrails was tricky. We solved this by designing an explicit human-in-the-loop gating mechanism that requires engineer authorization before critical scripts execute.
Accomplishments that we're proud of
Successfully integrating an asynchronous Python agent pipeline with a live-updating React frontend, achieving verified response metrics that slash incident MTTR significantly.
What we learned
We gained deep production-level experience in designing reliable agentic workflows, managing stateful asynchronous operations in FastAPI, and building clean operator dashboards in React.
What's next for ReplayOps
Expanding integrations with more cloud monitoring tools and adding advanced multi-agent predictive incident forecasting.
Built With
- api
- css
- fastapi
- javascript
- mongodb
- python
- react
- tailwind
Log in or sign up for Devpost to join the conversation.