FixPilot — Autonomous Production Incident Investigation & Resolution Agent

The Problem

When a production incident occurs, engineers often have to manually jump between logs, API responses, deployment history, source-code changes, and regression tests to determine what actually went wrong.

The difficult part isn't simply finding information. Engineers must decide:

  • What should I investigate first?
  • Which evidence is relevant?
  • What hypothesis should I test next?
  • Is the suspected root cause actually validated?
  • What action should be taken?
  • Which actions are safe to automate?

This investigation process is repetitive, time-consuming, and prone to incomplete evidence gathering.

The Solution

FixPilot is an autonomous production-incident investigation agent designed for software engineers and SREs.

Instead of following a fixed investigation script, FixPilot gives an AI agent access to a set of investigation tools. The agent decides which tools to call, provides their arguments, observes the results, and determines what to investigate next.

For a simulated incident such as:

Incident: INC-1001
Service: product-service
Error: KeyError: price

FixPilot can investigate logs, inspect API responses, examine deployment history, compare service versions, and run regression tests before producing a structured incident report.

Potentially dangerous actions such as deployment rollback are protected by a deterministic human-approval boundary.

What Makes It Agentic?

FixPilot is intentionally designed as an agent, rather than a deterministic workflow with an LLM attached to it.

The agent is responsible for:

  • Selecting investigation tools
  • Choosing tool arguments
  • Interpreting tool results
  • Forming and revising hypotheses
  • Deciding whether additional evidence is required
  • Producing the final investigation report

The surrounding application remains deterministic for execution and safety.

This creates a clear separation:

AI decides what to investigate. Software executes the investigation. Humans control risky production changes.

Investigation Tools

FixPilot currently provides tools for:

Log Search

Search simulated production logs for service-specific errors.

Deployment History

Inspect recent deployments and identify versions associated with an incident.

API Inspection

Inspect simulated API responses to determine whether unexpected input or response data may contribute to an incident.

Version Comparison

Compare the implementation of two service versions to identify potentially relevant code changes.

Regression Testing

Run a regression test against the suspected failure to validate whether the failure can actually be reproduced.

Rollback

Propose and simulate a deployment rollback.

Rollback is classified as a risky action and is blocked unless human approval is explicitly provided.

Structured Investigation Report

Instead of returning an unstructured answer, FixPilot produces a structured report containing:

  • Incident ID
  • Observed evidence
  • Inferences
  • Unresolved hypotheses
  • Root cause
  • Recommendations

This distinction is important because an AI agent should not present an assumption as a verified production fact.

Human-in-the-Loop Safety

FixPilot separates investigation from production-changing actions.

The agent can autonomously investigate an incident and recommend a rollback, but the rollback tool enforces an explicit approval requirement.

Investigation
     ↓
Evidence
     ↓
Hypothesis
     ↓
Recommendation
     ↓
Human Approval
     ↓
Production-changing Action

This allows autonomous investigation without giving the model unrestricted authority over production systems.

Architecture

React
  ↓
FastAPI
  ↓
Strands Agent
  ↓
Qwen3 4B / Ollama
  ↓
Autonomous Agent Loop
  ↓
Investigation Tools
  ├── Logs
  ├── Deployments
  ├── API
  ├── Versions
  ├── Regression Tests
  └── Rollback
          ↓
    Human Approval

The Strands agent sits at the center of the system and coordinates the investigation.

Technology

  • Python
  • Strands Agents SDK
  • Qwen3 4B
  • Ollama
  • FastAPI
  • Pydantic
  • React
  • Vite

The current prototype runs Qwen3 locally through Ollama and uses simulated production data. This keeps the demonstration safe and reproducible while allowing the complete agent/tool/safety architecture to be demonstrated.

Why It Matters

For engineers responding to production incidents, the valuable outcome is not another chatbot that explains an error.

The valuable outcome is an agent that can:

Investigate → Gather evidence → Test hypotheses → Validate failures → Recommend action

while keeping engineers in control of consequential production changes.

FixPilot explores that boundary between autonomous investigation and human-controlled production operations.

Current Status

FixPilot is a hackathon prototype demonstrating the core autonomous investigation architecture.

Implemented capabilities include:

  • Strands-based agent
  • Local Qwen3 model integration
  • Autonomous tool selection
  • Production log investigation
  • Deployment investigation
  • API inspection
  • Version comparison
  • Regression testing
  • Structured investigation reports
  • Human approval safety boundary
  • FastAPI investigation endpoint
  • React incident submission interface

The production environment is currently simulated.

Future versions can replace the simulated tools with integrations for real observability platforms, deployment systems, source-control systems, testing infrastructure, and production approval workflows.

Built for the Agents for Humans Hackathon

Track: Professional Agents

FixPilot is designed for software engineers and SREs who spend significant time investigating production incidents.

The project demonstrates how Strands Agents can turn a traditionally manual, multi-system investigation into an autonomous workflow while preserving a human approval boundary for risky actions.

Built With

  • agents
  • ai
  • calling
  • devops
  • fastapi
  • generative
  • llm
  • ollama
  • pydantic
  • python
  • qwen3
  • react
  • sdk
  • strands
  • tool
  • vite
Share this project:

Updates

Submission history