Inspiration

We started with a simple problem: production incidents take a lot of time to investigate.

When something breaks, an engineer usually has to:

  • Check the logs
  • Go through the stack trace
  • Find the relevant code
  • Figure out what caused the issue
  • Write a fix
  • Run tests
  • Create a pull request

We realized that a large part of this process is repetitive.

So we asked ourselves:

What if an AI agent could handle the repetitive investigation and fixing, while the engineer still makes the final decision?

That became the idea behind DevOpsSentinel.

What it does

DevOpsSentinel acts like an AI teammate for DevOps and SRE teams.

It takes an incident and works through the resolution process:

Jira → Logs → Investigation → Fix → QA Review → Tests → GitHub PR

The agent can:

  • Understand the production incident
  • Analyze the logs and stack trace
  • Look through the relevant source code
  • Find the likely root cause
  • Generate a code fix
  • Ask a separate QA/Safety Agent to review the fix
  • Improve the fix if the review fails
  • Run tests to make sure the change works
  • Create a GitHub pull request

The engineer is still in control. The system prepares the fix, but the final merge requires human approval.

How we built it

We used Strands Agents SDK to build the agent workflow and Amazon Bedrock for the AI reasoning.

The main technologies we connected were:

  • Amazon Bedrock — AI reasoning
  • Strands Agents SDK — agent orchestration
  • Jira — incident input
  • GitHub API — code and pull requests
  • Python — repository tools, patching, and testing
  • Streamlit — dashboard

We split the work between two agents.

Investigator Agent

This agent focuses on understanding the problem.

It looks at the incident, logs, stack trace, and source code before proposing a fix.

QA/Safety Agent

We didn't want the same agent to create a fix and immediately approve its own work.

So we added a second agent to review the proposed change.

It checks whether the fix actually addresses the issue, preserves existing behavior, and avoids unsafe changes.

Our demo

For the demo, we created a payment-service incident where the application fails with:

KeyError: 'billing_address'

When the incident is triggered, DevOpsSentinel receives the Jira ticket and production log.

It then:

Investigates → Finds the root cause → Generates the fix → Reviews it → Runs tests → Creates a GitHub PR

This was important to us because we wanted to demonstrate more than just an AI-generated explanation.

We wanted to show the complete journey from a real incident to a verified code change.

Challenges we faced

One challenge was getting the agent to look at the right context.

A log message alone doesn't always explain why something failed. We had to connect the incident with the stack trace and the relevant source code.

Another challenge was safety.

We didn't want an AI-generated patch to be applied just because it looked correct. The QA/Safety Agent and automated tests give us another layer of confidence before the change reaches the engineer.

We also spent time deciding where the AI should stop.

For us, the right approach was automation with human control. The agent can do the investigation and prepare the change, but the engineer gets the final say.

What we're proud of

The part we're most proud of is that DevOpsSentinel doesn't stop at telling us what went wrong.

It can take the incident through the actual development workflow:

Incident → Root Cause → Code Fix → Review → Tests → Pull Request

We also liked the idea of having one agent solve the problem and another agent challenge the solution.

That made the system feel much closer to having an AI teammate and a second reviewer working together.

What we learned

One of the biggest things we learned is that building an AI agent is not just about writing a prompt.

The agent needs:

  • Good context
  • Access to real tools
  • A structured workflow
  • Validation
  • Clear boundaries

We also learned that AI doesn't have to replace the engineer to be useful.

Even if the engineer still makes the final decision, removing the repetitive investigation work can make incident response much faster.

What's next

This is just the beginning for DevOpsSentinel.

We'd like to support more types of production failures and connect it with more observability and DevOps tools.

We also want to improve the agent's root-cause analysis and make patch verification stronger.

Our long-term idea is simple:

Let the AI handle the repetitive work. Let the engineer make the important decisions.

Built With

  • agentic-ai
  • ai-agents
  • amazon-bedrock
  • amazon-web-services
  • aws-cloudwatch
  • devops
  • fastapi
  • github-api
  • jira
  • multi-agent-systems
  • python
  • strands-agents-sdk
  • streamlit
Share this project:

Updates

Submission history