Inspiration

Inspiration

Production incidents can happen at any time, especially during high-traffic situations. When an outage happens at 2 AM, engineers often have to search through old post-mortems, logs, commands, and documentation to understand what happened before.

We wanted to build an AI SRE agent that does more than just analyze the current incident. Our idea was to give the agent memory of previous incidents, so it can learn from past failures and help engineers respond faster and more intelligently.

What it does

PostMortem AI is a self-healing Incident SRE Agent that helps engineers investigate and resolve production incidents.

It analyzes an incident, identifies possible root causes, searches its memory of previous incidents and post-mortems, and recommends solutions based on what worked in the past.

For example, if a PostgreSQL connection pool becomes exhausted during a sudden traffic spike, the agent can retrieve similar historical incidents, identify the previous root cause, and recommend the commands or configuration changes that helped resolve it.

The system turns previous incident knowledge into actionable recommendations for future incidents.

How we built it

We built the system around an AI agent and a persistent incident-memory layer.

The main workflow is:

Incident → Analysis → Memory Retrieval → Root Cause → Recommended Action → Resolution → New Memory

We use historical post-mortems, root causes, environment configurations, troubleshooting commands, and successful resolutions as the agent's memory.

The agent retrieves relevant past incidents and uses that information together with the current incident context to generate a response and recommended remediation steps.

Challenges we ran into

One of our biggest challenges was designing the memory so that the AI could retrieve the right historical incident instead of simply returning similar-looking information.

Another challenge was converting unstructured post-mortems and troubleshooting notes into useful information that an AI agent could understand and reuse.

We also had to design the system so that the AI's recommendations were based on previous successful resolutions rather than generic troubleshooting suggestions.

Accomplishments that we're proud of

We are proud of creating an AI-powered SRE concept that focuses on learning from previous incidents instead of starting from zero every time.

The biggest accomplishment is making persistent memory the core of the incident-response workflow. The system can connect a current production problem with historical incidents, root causes, configurations, and successful fixes.

We also designed the project around a realistic production scenario involving PostgreSQL connection-pool exhaustion during a sudden checkout traffic spike on AWS ECS.

What we learned

We learned that an effective AI agent needs more than a powerful model. Context, memory, retrieval, and reliable historical information are equally important.

We also learned how incident-response workflows can be transformed into reusable knowledge. Every resolved incident can become a valuable resource for handling future incidents.

Most importantly, we learned how persistent AI memory can help engineering teams move from reactive troubleshooting toward faster, knowledge-driven incident resolution.

What it does

How we built it

Challenges we ran into

Accomplishments that we're proud of

What we learned

What's next for PRO Caps

Built With

Share this project:

Updates

Submission history