Multi-agent incident response system that automates production debugging. When alerts fire, three AI agents work in parallel to analyze logs, metrics, and deployments — then synthesize findings into a structured incident report with root cause analysis.

Problem

When production breaks, engineers waste 30+ minutes context-switching between:

  • Log dashboards (Cloud Logging, ELK)
  • Metrics views (Cloud Monitoring, Grafana)
  • Deployment history (Cloud Run, GitHub)
  • Slack threads and runbooks

This is meant to be configured once with connectors to all yyour services, especially google services and agregating logs in one place then enabling. Gemini based bug tracking.

Built With

Share this project:

Updates