-
-
Hero/Title Image
-
quickstart
-
Scenario 1: Ad-Ranking Crash
-
Scenario 5: PII Exposure
-
Scenario 2: Timestamp Corruption
-
Scenario 3: Connection Pool Leak
-
metric
-
Architecture Diagram (3-layer system)
-
Scenario 4: Cascading Failure + Red Herrings
-
Agent Workflow Diagram (decision loop)
-
impact
-
rewards
-
Agent loop
-
comparsion
-
tools
Inspiration
Production incidents are difficult to diagnose because the visible failure is often not the real cause. A service may show high latency, failed requests, or database errors while the actual problem is an upstream data bug, a resource leak, or a faulty deployment.
We wanted to explore whether an AI agent could learn to investigate incidents the way an experienced Site Reliability Engineer does: by checking logs, inspecting metrics, tracing dependencies, making a focused code change, running the correct tests, and documenting the root cause.
This led us to build META-SRE, an open-source training and evaluation environment for Agentic AI systems focused on production incident response.
What the project does
META-SRE places an AI agent inside a simulated microservice environment containing three services:
ad_rankingcapi_pipelinewhatsapp_sync
The agent receives an incident description, system metrics, alerts, terminal output, service dependencies, and a limited step budget. It can then use SRE tools to investigate and resolve the incident.
The available tools include:
- Viewing service files
- Applying surgical line-by-line code edits
- Reading service logs
- Querying metric history
- Checking dependencies
- Running unit, integration, load, and security tests
- Inspecting simulated Git blame
- Rolling back database migrations
- Asking a senior SRE for a hint
- Writing a final incident report
The agent is not evaluated only on whether it changes the code. It is also evaluated on:
- Whether it finds the true root cause
- Whether it avoids fixing downstream symptoms
- Whether the tests pass
- Whether the incident is resolved within the SLA budget
- Whether the final incident report is accurate
- Whether the solution avoids regressions
How we built it
META-SRE is implemented in Python using FastAPI, Pydantic, and an OpenEnv-compatible API.
The system is divided into several components:
1. Virtual incident environment
A virtual in-memory file system stores the simulated service codebases. Each episode starts from a clean buggy snapshot, allowing the same incident to be reproduced consistently.
The environment supports:
- File inspection
- Line-level edits
- Edit history
- Git-diff generation
- Simulated Git blame
2. Observability engine
The metrics engine generates service telemetry for every step of an episode, including:
- CPU usage
- Memory usage
- Error rate
- P99 latency
- Request queue depth
- Deployment timestamps
- Service health status
It also produces alerts and simulated terminal output such as stack traces, database errors, and security warnings.
3. Agent tool layer
The tool dispatcher exposes structured JSON tool specifications to the agent. Every action produces an output, reward value, and updated observation.
This creates a controlled decision-making loop:
- Observe the incident
- Select an investigation tool
- Analyze the result
- Modify the simulated environment
- Run validation tests
- Submit an incident report
4. Task grader and reward system
The grader checks whether the agent actually fixed the underlying issue. The reward system combines:
- Step penalties for inefficient actions
- Penalties for syntax errors and incorrect rollbacks
- Progress rewards for finding the correct service and file
- Test-passing rewards
- Incident report accuracy
- SLA and no-regression bonuses
This encourages the agent to solve incidents efficiently rather than randomly editing files.
Incident scenarios
The current environment contains five progressive scenarios:
Ad-ranking crash
A dictionary is incorrectly treated as an object, causing anAttributeError.Silent timestamp corruption
A timestamp normalization threshold corrupts event times and causes incorrect advertising attribution.Database connection leak
Async database connections are not released, exhausting the connection pool under peak load.Cascading failure caused by a migration
A circular foreign-key relationship causes multiple services to fail simultaneously, while downstream alerts act as red herrings.Production PII exposure
Debug mode causes raw user data to be returned in an API response. Unit tests pass, but the security suite detects the vulnerability.
What we learned
While building META-SRE, we learned that an effective agent environment needs more than a chatbot interface. The environment must provide:
- Structured observations
- Actions with clear side effects
- Realistic but reproducible failures
- Feedback that rewards useful progress
- Evaluation criteria that measure reasoning, not just final output
We also learned that different incidents require different validation strategies. Unit tests are not sufficient for load-dependent failures or security vulnerabilities. Agents must learn to choose the appropriate test suite based on the evidence available.
Challenges we faced
The main challenges were:
- Designing realistic incidents without requiring real infrastructure
- Simulating service dependencies and failure propagation
- Making silent data corruption distinguishable from normal service behavior
- Preventing agents from modifying test files to achieve artificial success
- Balancing exploration with a limited step and reward budget
- Evaluating incident reports in a consistent and explainable way
- Keeping each episode isolated and reproducible
We addressed these challenges with an in-memory virtual file system, deterministic task snapshots, structured task graders, contextual logs, simulated metrics, and a reward manager.
Try the project
META-SRE exposes a FastAPI interface with endpoints for:
- Starting an incident
- Executing agent actions
- Viewing the current state
- Listing tools and tasks
- Viewing the final episode grade
The project can be run locally with Docker or directly using Uvicorn. The interactive API documentation is available through FastAPI Swagger UI.
Future plans
Our next steps are to:
- Add more incident categories and service topologies
- Add multi-agent collaboration between SRE agents
- Improve the training-data generation pipeline
- Add persistent experiment tracking
- Connect the environment to open-source language models
- Add a dashboard for comparing agent performance
- Introduce optional blockchain-based verification for training trajectories and achievement records
META-SRE is intended to become a reusable open-source benchmark and training environment for AI agents that must reason about reliability, security, and production operations.
Log in or sign up for Devpost to join the conversation.