Inspiration
Incident response becomes difficult quickly under pressure.
Engineers need to collect evidence, understand what changed, decide whether a remediation is safe, and recover the service without making the incident worse.
AI can help with investigation and planning, but giving an LLM direct control over production introduces a new risk: the system may be able to reason about a change and execute it at the same time.
IncidentPilot was built around a simple principle:
AI investigates and recommends. Humans authorize. Deterministic code executes and verifies.
The goal is to demonstrate how agentic AI can support SRE incident response while keeping remediation under explicit human control.
What it does
IncidentPilot is an AI-assisted incident investigation and recovery workflow for SRE engineers.
The demo starts with a synthetic ECS-like production incident:
payments-apiversionv43has just been deployed- the task is
STOPPED - memory usage reaches
99% - the container exits with code
137 - the error rate rises to
12% - the health endpoint returns HTTP
503
From that point, IncidentPilot orchestrates exactly three specialized AWS Strands agents.
Each agent has a deliberately limited responsibility and no direct execution authority.
1. Incident Investigator
The Investigator uses read-only tools to inspect:
- deployment history
- task status
- memory metrics
- error metrics
- application logs
- health status
It correlates the evidence and produces a structured investigation report containing:
- likely root cause
- confidence
- supporting evidence
- remaining unknowns
- deployment context for the next agent
The Investigator can analyze evidence, but it cannot modify the service.
2. Remediation Planner
The Planner receives the investigation and proposes a remediation plan.
For the demo scenario, it recommends rolling back from v43 to the previously healthy v42 release.
The plan contains:
- the proposed action
- the reason for the change
- the expected outcome
- risks and trade-offs
- verification criteria
This remains only a proposal.
The Planner cannot execute the rollback and cannot authorize it.
3. Safety Reviewer
The Safety Reviewer independently evaluates the investigation and remediation plan.
It considers:
- evidence quality
- uncertainty
- risk level
- blast radius
- consistency between the evidence and the proposed action
The result can be APPROVED, REJECTED, or require further human review.
However, even an APPROVED result is explicitly advisory.
A Safety Reviewer approval does not authorize execution.
Human-controlled remediation
Human approval is a separate control boundary.
The engineer must explicitly approve or reject the proposed remediation.
Only a valid human approval can unlock execution.
Even then, approval itself does not modify the environment.
The engineer must separately choose to execute the approved action.
This separation is intentional:
AI recommendation ≠ human authorization ≠ execution
Deterministic execution and verification
The LLM agents never execute the rollback.
A deterministic Python executor applies the approved synthetic rollback from v43 to v42.
After execution, a separate deterministic verifier checks recovery.
The expected recovered state is:
- active version:
v42 - task status:
RUNNING - memory:
48% - error rate:
0.3% - health: HTTP
200
The incident is not marked as resolved simply because the rollback executed.
Only after all verification checks pass does IncidentPilot transition the incident to:
RESOLVED
This creates a clear separation between reasoning, authorization, execution, and verification.
How we built it
IncidentPilot is built with:
- AWS Strands Agents for the three specialized agents
- Amazon Bedrock for model inference
- Claude Sonnet 4.6 through an AWS Bedrock inference profile for the validated live demo
- FastAPI for the workflow API
- Pydantic for strongly validated structured agent outputs
- plain HTML, CSS, and JavaScript for the demo control room
- deterministic Python components for approval, execution, and verification
The application intentionally separates agent reasoning from environment mutation.
The default API is provider-safe and can load without creating a model connection.
Real Bedrock inference is enabled only through an explicit runtime configuration.
For testing, the Strands workflow can run with scripted offline models, so the test suite does not require paid inference or AWS credentials.
The demo uses synthetic and repeatable infrastructure evidence and performs no real AWS infrastructure writes.
Architecture
The IncidentPilot workflow is:
Incident detected → Investigator → Planner → Safety Reviewer → Human Approval → Deterministic Execution → Deterministic Verification → RESOLVED
The architecture deliberately separates three areas:
AI reasoning
The three AWS Strands agents:
- investigate
- propose
- review
They never mutate the environment.
Human control
The engineer explicitly decides whether remediation is authorized.
Deterministic safety
Validated application code performs execution and verification.
This makes the human control boundary visible both in the system architecture and in the user interface.
Challenges we faced
Deciding where AI autonomy should stop
One of the most important design challenges was defining the exact boundary between AI reasoning and execution.
It was not enough to tell the Safety Reviewer in a prompt that it should not execute changes.
We needed the application itself to guarantee that an APPROVED review could never accidentally become execution authorization.
Human approval is therefore modeled as a separate validated record, and the executor independently checks for both:
- an approved safety review
- an explicit approved human decision
Without both, execution fails closed.
Keeping evidence separate from inference
Another challenge was ensuring that the agents did not overstate what the evidence proves.
For example, exit code 137 indicates a SIGKILL, but by itself it does not prove an out-of-memory event.
The Investigator and Safety Reviewer were refined so memory exhaustion is inferred only when corroborated by additional signals such as:
- memory utilization
- deployment timing
- task termination
- error-rate increase
- service health degradation
This helps keep observed facts separate from inferred root cause.
Testing agents without continuous paid inference
We also wanted the project to remain easy to test and reproduce.
The workflow therefore supports scripted Strands models for offline tests while keeping real Amazon Bedrock inference as an explicit opt-in runtime.
This allowed us to validate the workflow repeatedly without making every test dependent on network access or paid model calls.
What we learned
IncidentPilot reinforced an important idea:
Agentic systems do not need unrestricted execution privileges to be useful.
The three-agent design worked well because each agent has a narrow responsibility:
- investigation
- remediation planning
- safety review
This made the workflow easier to reason about, test, and audit.
We also learned that safety should not rely only on prompts.
Prompts help define agent behavior, but the most important controls are enforced outside the LLM:
- human approval
- execution authorization
- supported remediation actions
- deterministic verification
- fail-closed workflow rules
Structured outputs were also valuable.
Using validated Pydantic models between stages allowed every step to verify the output of the previous one before continuing.
What's next
The current version focuses on a safe and repeatable synthetic incident-response workflow.
Future work could include:
- connecting read-only evidence tools to real AWS services such as ECS and CloudWatch
- supporting additional incident scenarios and remediation strategies
- adding persistent incident history and audit records
- introducing richer verification policies
- supporting additional rollback and recovery strategies
- evaluating AWS AgentCore for production-grade agent runtime capabilities
- extending the control-room UI for multi-incident workflows
The core principle would remain unchanged:
AI investigates and recommends. Humans authorize. Deterministic code executes and verifies.
Built With
- agents
- ai
- amazon
- amazon-web-services
- bedrock
- css
- devops
- fastapi
- generative
- html
- human-in-the-loop
- incident
- javascript
- python
- sre
- strands
Log in or sign up for Devpost to join the conversation.