Inspiration

Incident response becomes difficult quickly under pressure.

Engineers need to collect evidence, understand what changed, decide whether a remediation is safe, and recover the service without making the incident worse.

AI can help with investigation and planning, but giving an LLM direct control over production introduces a new risk: the system may be able to reason about a change and execute it at the same time.

IncidentPilot was built around a simple principle:

AI investigates and recommends. Humans authorize. Deterministic code executes and verifies.

The goal is to demonstrate how agentic AI can support SRE incident response while keeping remediation under explicit human control.


What it does

IncidentPilot is an AI-assisted incident investigation and recovery workflow for SRE engineers.

The demo starts with a synthetic ECS-like production incident:

  • payments-api version v43 has just been deployed
  • the task is STOPPED
  • memory usage reaches 99%
  • the container exits with code 137
  • the error rate rises to 12%
  • the health endpoint returns HTTP 503

From that point, IncidentPilot orchestrates exactly three specialized AWS Strands agents.

Each agent has a deliberately limited responsibility and no direct execution authority.

1. Incident Investigator

The Investigator uses read-only tools to inspect:

  • deployment history
  • task status
  • memory metrics
  • error metrics
  • application logs
  • health status

It correlates the evidence and produces a structured investigation report containing:

  • likely root cause
  • confidence
  • supporting evidence
  • remaining unknowns
  • deployment context for the next agent

The Investigator can analyze evidence, but it cannot modify the service.

2. Remediation Planner

The Planner receives the investigation and proposes a remediation plan.

For the demo scenario, it recommends rolling back from v43 to the previously healthy v42 release.

The plan contains:

  • the proposed action
  • the reason for the change
  • the expected outcome
  • risks and trade-offs
  • verification criteria

This remains only a proposal.

The Planner cannot execute the rollback and cannot authorize it.

3. Safety Reviewer

The Safety Reviewer independently evaluates the investigation and remediation plan.

It considers:

  • evidence quality
  • uncertainty
  • risk level
  • blast radius
  • consistency between the evidence and the proposed action

The result can be APPROVED, REJECTED, or require further human review.

However, even an APPROVED result is explicitly advisory.

A Safety Reviewer approval does not authorize execution.


Human-controlled remediation

Human approval is a separate control boundary.

The engineer must explicitly approve or reject the proposed remediation.

Only a valid human approval can unlock execution.

Even then, approval itself does not modify the environment.

The engineer must separately choose to execute the approved action.

This separation is intentional:

AI recommendation ≠ human authorization ≠ execution


Deterministic execution and verification

The LLM agents never execute the rollback.

A deterministic Python executor applies the approved synthetic rollback from v43 to v42.

After execution, a separate deterministic verifier checks recovery.

The expected recovered state is:

  • active version: v42
  • task status: RUNNING
  • memory: 48%
  • error rate: 0.3%
  • health: HTTP 200

The incident is not marked as resolved simply because the rollback executed.

Only after all verification checks pass does IncidentPilot transition the incident to:

RESOLVED

This creates a clear separation between reasoning, authorization, execution, and verification.


How we built it

IncidentPilot is built with:

  • AWS Strands Agents for the three specialized agents
  • Amazon Bedrock for model inference
  • Claude Sonnet 4.6 through an AWS Bedrock inference profile for the validated live demo
  • FastAPI for the workflow API
  • Pydantic for strongly validated structured agent outputs
  • plain HTML, CSS, and JavaScript for the demo control room
  • deterministic Python components for approval, execution, and verification

The application intentionally separates agent reasoning from environment mutation.

The default API is provider-safe and can load without creating a model connection.

Real Bedrock inference is enabled only through an explicit runtime configuration.

For testing, the Strands workflow can run with scripted offline models, so the test suite does not require paid inference or AWS credentials.

The demo uses synthetic and repeatable infrastructure evidence and performs no real AWS infrastructure writes.


Architecture

The IncidentPilot workflow is:

Incident detected → Investigator → Planner → Safety Reviewer → Human Approval → Deterministic Execution → Deterministic Verification → RESOLVED

The architecture deliberately separates three areas:

AI reasoning

The three AWS Strands agents:

  • investigate
  • propose
  • review

They never mutate the environment.

Human control

The engineer explicitly decides whether remediation is authorized.

Deterministic safety

Validated application code performs execution and verification.

This makes the human control boundary visible both in the system architecture and in the user interface.


Challenges we faced

Deciding where AI autonomy should stop

One of the most important design challenges was defining the exact boundary between AI reasoning and execution.

It was not enough to tell the Safety Reviewer in a prompt that it should not execute changes.

We needed the application itself to guarantee that an APPROVED review could never accidentally become execution authorization.

Human approval is therefore modeled as a separate validated record, and the executor independently checks for both:

  • an approved safety review
  • an explicit approved human decision

Without both, execution fails closed.

Keeping evidence separate from inference

Another challenge was ensuring that the agents did not overstate what the evidence proves.

For example, exit code 137 indicates a SIGKILL, but by itself it does not prove an out-of-memory event.

The Investigator and Safety Reviewer were refined so memory exhaustion is inferred only when corroborated by additional signals such as:

  • memory utilization
  • deployment timing
  • task termination
  • error-rate increase
  • service health degradation

This helps keep observed facts separate from inferred root cause.

Testing agents without continuous paid inference

We also wanted the project to remain easy to test and reproduce.

The workflow therefore supports scripted Strands models for offline tests while keeping real Amazon Bedrock inference as an explicit opt-in runtime.

This allowed us to validate the workflow repeatedly without making every test dependent on network access or paid model calls.


What we learned

IncidentPilot reinforced an important idea:

Agentic systems do not need unrestricted execution privileges to be useful.

The three-agent design worked well because each agent has a narrow responsibility:

  • investigation
  • remediation planning
  • safety review

This made the workflow easier to reason about, test, and audit.

We also learned that safety should not rely only on prompts.

Prompts help define agent behavior, but the most important controls are enforced outside the LLM:

  • human approval
  • execution authorization
  • supported remediation actions
  • deterministic verification
  • fail-closed workflow rules

Structured outputs were also valuable.

Using validated Pydantic models between stages allowed every step to verify the output of the previous one before continuing.


What's next

The current version focuses on a safe and repeatable synthetic incident-response workflow.

Future work could include:

  • connecting read-only evidence tools to real AWS services such as ECS and CloudWatch
  • supporting additional incident scenarios and remediation strategies
  • adding persistent incident history and audit records
  • introducing richer verification policies
  • supporting additional rollback and recovery strategies
  • evaluating AWS AgentCore for production-grade agent runtime capabilities
  • extending the control-room UI for multi-incident workflows

The core principle would remain unchanged:

AI investigates and recommends. Humans authorize. Deterministic code executes and verifies.

Built With

Share this project:

Updates

Submission history