Inspiration

DevOps and cloud operations engineers often repeat the same steps whenever a service alarm fires: check the service status, inspect recent logs, determine whether a known remediation is safe, perform the action, verify recovery, and inform the responsible engineer.

Although these steps are necessary, they interrupt engineers and increase recovery time. A traditional watchdog can restart a process, but it cannot interpret evidence, explain its decision, distinguish a symptom from a root cause, or communicate uncertainty responsibly.

We built OpsPilot to automate this repetitive first-response process while keeping infrastructure actions controlled, transparent, and verifiable.

What it does

OpsPilot is an autonomous, event-driven incident-response agent for AWS workloads.

When the monitored demo-api service becomes unhealthy:

  1. A systemd timer checks the service every minute.
  2. The health check publishes a custom metric to Amazon CloudWatch.
  3. A CloudWatch alarm detects the unhealthy state.
  4. Amazon EventBridge routes the alarm event to AWS Lambda.
  5. Lambda uses AWS Systems Manager Run Command to launch OpsPilot on the designated EC2 instance.
  6. The Strands agent checks the current service status and collects recent system logs.
  7. An Amazon Bedrock-hosted model analyzes the evidence and determines the appropriate next step.
  8. A deterministic remediation policy allows only an approved restart of demo-api.
  9. OpsPilot verifies the internal /health endpoint and requires an HTTP 200 response before declaring recovery.
  10. Amazon SNS emails a structured incident report to the subscribed engineer.
  11. The restored health metric returns the CloudWatch alarm from ALARM to OK.

The report includes the initial service state, failure condition, relevant evidence, underlying root cause when supported, confidence level, remediation performed, verification result, and final incident status.

How we built it

We developed OpsPilot in Python using the Strands Agents SDK. The reasoning model is accessed through Amazon Bedrock Mantle using the Strands OpenAI Responses model integration.

The Strands agent has four focused tools:

  • get_service_status checks the current systemd service state.
  • get_recent_logs retrieves logs relevant to the current incident.
  • restart_service enforces the deterministic remediation policy.
  • verify_service_health checks the application’s internal health endpoint.

The surrounding event-driven architecture uses:

  • Amazon EC2
  • Amazon CloudWatch custom metrics and alarms
  • Amazon EventBridge
  • AWS Lambda
  • AWS Systems Manager Run Command
  • Amazon Bedrock
  • Amazon SNS
  • AWS IAM
  • Python, Bash, and systemd

Strands acts as the reasoning and tool-orchestration layer—not simply as a chatbot. It decides how to gather evidence, analyzes the latest operational state, selects an appropriate approved tool, verifies the result, and produces the final incident report.

We deliberately separated reasoning from authorization. The model can analyze evidence and request an action, but deterministic code decides whether that action is permitted.

For safety:

  • Only demo-api is approved for automated remediation.
  • Restart is the only approved action.
  • Unknown services are denied.
  • Evidence must be collected before remediation.
  • A stopped service is treated as a failure condition, not automatically as its root cause.
  • Unsupported root-cause conclusions are reported as unknown.
  • Recovery requires a successful HTTP 200 health check.
  • IAM permissions are limited to the required instance, command document, metric namespace, and SNS topic.
  • The application health port is not publicly exposed.
  • Deployment-specific identifiers and credentials are excluded from the public repository.

We also created a reusable CloudFormation template for the automation control plane and published a sanitized, MIT-licensed GitHub repository containing source code, tests, systemd files, IAM examples, deployment instructions, and the architecture diagram.

Challenges we ran into

The most important challenge was preventing the agent from confusing an observed failure condition with the underlying root cause.

During early testing, the agent could describe “service stopped” as the root cause. However, the logs only proved that the service was inactive; they did not explain why it stopped.

We refined the instructions so that the agent:

  • prioritizes the most recent evidence;
  • distinguishes symptoms from causes;
  • avoids relying on unrelated historical errors;
  • communicates uncertainty clearly;
  • does not assign high confidence without supporting evidence; and
  • explicitly reports an unknown root cause when the available evidence is insufficient.

We repeatedly tested the complete workflow by deliberately stopping the service and allowing OpsPilot to detect, investigate, remediate, verify, notify, and recover the CloudWatch alarm without a manual restart.

Another challenge was converting a functional prototype into a safer and more reusable project. We removed public access to the application port, restricted SSH access, tightened IAM permissions, moved deployment-specific identifiers into environment variables, sanitized the repository, and added infrastructure-as-code support.

Accomplishments that we're proud of

  • Built a real autonomous incident-response workflow instead of a prompt-only demonstration.
  • Connected Strands reasoning to real AWS operational tools.
  • Demonstrated repeated end-to-end recovery from a controlled service outage.
  • Enforced a deterministic service-and-action allowlist.
  • Preserved uncertainty instead of generating an unsupported root cause.
  • Verified recovery through the real application endpoint rather than trusting the restart command.
  • Delivered a complete incident report through Amazon SNS.
  • Returned the CloudWatch alarm from ALARM to OK.
  • Secured the implementation with scoped IAM permissions and restricted network access.
  • Published a reusable, sanitized, MIT-licensed repository with deployment instructions and CloudFormation.

In our controlled tests, OpsPilot detected and recovered the stopped service within approximately two to three minutes without manual investigation or restart.

What we learned

We learned that responsible agentic operations require two separate layers.

The reasoning layer interprets evidence and selects the appropriate tool. The deterministic layer controls what that tool is authorized to do. This makes the agent useful while preventing unrestricted infrastructure changes.

We also learned that successful command execution is not the same as successful recovery. A trustworthy incident-response agent must verify the real application endpoint before closing an incident.

Most importantly, we learned that meaningful autonomy does not require unnecessary architectural complexity. Every component in OpsPilot has a clear purpose: detection, event routing, evidence collection, reasoning, controlled remediation, verification, or notification.

What's next for OpsPilot

Future versions could add:

  • human approval for higher-risk remediation;
  • additional explicitly allowlisted services and runbooks;
  • richer notification and collaboration integrations;
  • persistent incident history and trend analysis;
  • evaluation datasets for measuring diagnostic accuracy;
  • enhanced observability and audit trails;
  • support for containerized and Kubernetes workloads; and
  • Amazon Bedrock AgentCore for managed deployment and production scaling.

The guiding principle will remain the same: the agent may reason broadly, but infrastructure actions must remain narrow, explicit, evidence-based, and verifiable.

Built With

Share this project:

Updates

Submission history