Inspiration

Production incidents don't wait for business hours, and small teams rarely have a dedicated on-call SRE. When an alarm fires at 3am, someone still has to wake up, grep through logs, check recent deploys, and piece together what happened — before they can even start fixing it. I wanted to see if an AI agent could do that first investigative pass entirely on its own, so a human's first message is already a root-cause summary instead of a blank slate.

What it does

IncidentPilot is an autonomous agent that responds to CloudWatch alarms with zero human input. When an alarm fires:

  1. A CloudWatch Alarm state change triggers an EventBridge rule
  2. EventBridge invokes a Lambda function, which builds an incident prompt from the raw alarm payload
  3. That Lambda calls an AI agent hosted on Amazon Bedrock AgentCore Runtime
  4. The agent autonomously decides which tools to use — it queries CloudWatch Logs Insights for recent errors, checks the GitHub repo for recent commits that might explain the issue, and reasons over both
  5. It writes a concise root-cause summary with a suggested fix, and posts it directly to Slack

No one triggers the investigation. No one reviews it before it's posted. The alarm firing is the only input.

How I built it

The agent itself is built with the Strands Agents SDK, with three custom tools (fetch_logs, get_recent_deploys, post_to_slack) exposed to the model. It's hosted on Bedrock AgentCore Runtime rather than running inline in a Lambda — AgentCore gives it isolated, managed execution as a first-class agent resource, while the trigger Lambda stays a thin dispatcher that just builds the prompt and calls InvokeAgentRuntime.

Infrastructure (the Lambda, its IAM role, and the EventBridge rule) is deployed with AWS SAM. All secrets (API keys, GitHub token, Slack webhook) are stored in SSM Parameter Store as SecureStrings and fetched at runtime — nothing is hardcoded or committed.

For the live demo, I built a small FastAPI app with a deliberately-broken /trigger-error endpoint, deployed as its own Lambda behind API Gateway, with a real CloudWatch metric filter and alarm watching its logs — so the whole pipeline can be triggered by anyone, live, with one click.

Challenges I ran into

  • Bedrock model access got restricted mid-build by an account-level hold on a brand-new AWS account, which also briefly blocked Lambda entirely. I diagnosed it down to an account-level issue (not IAM, not an SCP) and worked with AWS Support to get it lifted — then, to avoid being blocked by it again, added support for calling Claude directly via the Anthropic API as a fallback model provider alongside Bedrock.
  • AgentCore's tooling changed under me — the CLI I was using got renamed mid-project (launch → deploy), and environment variables didn't propagate the way I expected on the first few tries, which took real debugging via CloudWatch Logs on the AgentCore Runtime side to track down.
  • AWS App Runner turned out to be closed to new customers as of earlier this year, so the demo trigger app had to be re-architected onto Lambda + API Gateway instead, using Mangum to adapt the existing FastAPI app with no code changes.

Accomplishments that I'm proud of

Getting a genuinely autonomous, end-to-end pipeline working — where a real error in a real service trips a real CloudWatch alarm and the entire investigation and Slack post happens with no manual trigger at any point. It's also fully reproducible from the README with no hardcoded secrets anywhere in the repo.

What I learned

A lot about how CloudWatch Alarms, EventBridge, and Lambda actually wire together at the event-shape level, and how Bedrock AgentCore Runtime differs from just running an agent inside a Lambda — particularly around execution roles, environment variable propagation, and the toolkit's still-evolving CLI.

What's next for IncidentPilot

  • Support for more notification channels (PagerDuty, email) alongside Slack
  • A memory layer so the agent can recognize recurring incidents and reference past fixes
  • Expanding the toolset to include a "safe rollback" action the agent can propose (and, with approval, trigger)

Built With

  • amazon-bedrock-agentcore
  • amazon-cloudwatch
  • amazon-eventbridge
  • anthropic-api
  • aws-api-gateway
  • aws-lambda
  • aws-sam
  • aws-ssm-parameter-store
  • claude
  • fastapi
  • github
  • python
  • slack-api
  • strands-agents-sdk
Share this project:

Updates

Submission history