Inspiration

I run a Linux/Docker homelab with services such as Jellyfin, Sonarr, Radarr, qBittorrent, Gluetun, Caddy, Prowlarr, Bazarr, SABnzbd, and Netdata.

As the number of services grew, I noticed that troubleshooting became much harder than simply checking whether a container was running. A visible problem in one service could actually be caused by another service several dependencies away.

For example, Sonarr might be running normally while torrent downloads fail because qBittorrent has lost connectivity through Gluetun. Jellyfin itself might be healthy while remote access is broken because the reverse proxy is unavailable.

Traditional monitoring is good at telling me what is down, but I wanted something that could investigate why it is broken, understand dependencies, propose a safe recovery plan, and then verify that the original problem was actually fixed.

That became HomelabOps.

What it does

HomelabOps is an AI-assisted DevOps agent for Linux and Docker homelabs.

A user can describe an infrastructure problem in natural language, for example:

“Sonarr torrent downloads are not working. Investigate and remediate it if appropriate.”

The agent then gathers live evidence using constrained infrastructure tools. It can inspect:

  • Docker container state and health
  • container lifecycle events and restart history
  • logs
  • system load, memory, swap, and disk usage
  • local HTTP service health
  • public endpoints
  • service dependencies
  • workflow dependencies
  • service-to-service connectivity

HomelabOps does not treat every dependency failure as a complete outage. It reasons about individual capabilities such as torrent_downloads, usenet_downloads, subtitles, and remote access.

It then records a structured incident containing the diagnosis, severity, confidence, root cause, affected service, and affected capability.

If remediation is appropriate, the agent can create a structured remediation proposal or multi-step recovery plan.

Human-in-the-loop safety

I wanted the agent to be useful without giving an LLM unrestricted control over my infrastructure.

The model therefore does not execute arbitrary shell commands.

Supported state-changing operations are implemented outside the model and limited to validated actions such as:

  • starting a known container
  • restarting a known container

Before one of those actions executes, HomelabOps requires explicit human approval.

Configuration defects are even more restricted. The agent can diagnose the issue and propose the intended configuration change, but configuration proposals cannot be automatically executed.

After an approved remediation runs, HomelabOps still does not assume that the problem is solved.

It performs post-action verification, which can include:

  1. container recovery
  2. Docker health
  3. application health
  4. public endpoint availability
  5. service-to-service connectivity
  6. the original affected workflow capability

Only after the relevant checks succeed can an incident be marked resolved.

How I built it

The core agent is built with the Strands Agents SDK and uses Amazon Bedrock as the model provider.

The agent interacts with the homelab through a collection of purpose-built tools rather than unrestricted command execution.

These tools cover areas such as:

  • Docker status
  • container details
  • lifecycle events
  • logs
  • resource monitoring
  • dependency inspection
  • workflow health
  • connectivity checks
  • service health
  • external endpoint verification
  • incident classification
  • remediation planning

The homelab itself is accessed over SSH using constrained Python tooling.

For the application layer, I built a FastAPI backend and a React + TypeScript + Vite frontend.

The web console contains five main views:

Overview

A live view of Docker services and host resources.

AI Agent

The natural-language interface for starting infrastructure investigations.

Incidents

A structured list of previous investigations with severity, confidence, root cause, affected capability, and a chronological event timeline.

Remediation

Shows pending actions, previous proposals, remediation plans, and verification results.

History

A searchable audit trail of agent investigations, approvals, executions, and verification events.

There is also a CLI interface. Both the CLI and web application share the same remediation execution and verification pipeline so that safety-critical logic is not duplicated.

A real example

One of the most interesting test cases involved Sonarr and qBittorrent.

Sonarr was running, but its torrent-download capability was broken.

HomelabOps investigated the workflow dependencies and found that Sonarr could not reach qBittorrent. It then inspected the relevant Docker lifecycle events and discovered that Gluetun had restarted while qBittorrent had remained running.

The diagnosis was that qBittorrent was still attached to a stale Gluetun network namespace.

HomelabOps created a recovery plan:

  1. verify Gluetun is healthy
  2. restart qBittorrent with human approval
  3. verify Sonarr → qBittorrent connectivity
  4. verify Sonarr's torrent_downloads capability

After approval, the remediation executed and each verification step passed before the incident was marked resolved.

Another test container intentionally had a health check configured as:

CMD-SHELL exit 1

Instead of repeatedly restarting the container, HomelabOps recognized that the health check was deterministically broken while the underlying nginx application was still operational. It proposed a configuration change instead of an ineffective restart.

That distinction was one of the behaviors I most wanted from the project.

Challenges

One of the biggest challenges was preventing the model from making reasonable-sounding but unsupported assumptions.

For example:

  • a running container is not necessarily healthy
  • container uptime is not proof that it has never restarted
  • an exit code of 0 does not prove that a human intentionally stopped a service
  • a running dependency does not prove that another service can reach it
  • recovering a container does not prove that the original workflow has recovered

I addressed this by making the system prompt strongly evidence-driven and by adding dedicated tools for lifecycle events, health checks, dependencies, connectivity, and workflow verification.

Another challenge was making the CLI and web interface behave identically. Initially, some remediation and incident-state behavior existed only in the CLI. I refactored the safety-critical execution logic into a shared remediation_execution.py module so both interfaces use the same approval, execution, and verification path.

The web interface also introduced a new security consideration because the FastAPI backend can trigger approved infrastructure actions. The current version is deliberately designed as a localhost-only management console rather than pretending that CORS is authentication.

What I learned

This project taught me that building an infrastructure agent is much more than connecting an LLM to shell commands.

The most important work was defining:

  • what evidence the agent is allowed to trust
  • what the model can and cannot execute
  • how dependencies should be represented
  • how partial degradation differs from an outage
  • when human approval is required
  • what evidence is sufficient to call an incident resolved

I also learned how important structured state becomes once an agent performs multi-step work. Incident results, remediation plans, audit events, and verification results make the system much easier to understand and debug than relying only on conversational output.

Accomplishments

I am particularly happy that HomelabOps now includes:

  • a working Strands/Bedrock agent
  • dependency-aware root-cause analysis
  • capability-level workflow diagnosis
  • human-approved remediation
  • multi-step remediation plans
  • post-action recovery verification
  • a React operations console
  • incident timelines
  • remediation tracking
  • searchable audit history
  • a shared CLI/web execution pipeline
  • automated safety tests

The current automated test suite contains 64 passing tests, including API tests that ensure automated testing cannot accidentally perform live SSH remediation.

What's next

The current version runs as a localhost-only operations console.

Future improvements could include:

  • authenticated remote access
  • Amazon Bedrock AgentCore deployment
  • additional infrastructure providers
  • more service integrations
  • alert-driven automatic investigation
  • richer metrics and observability
  • persistent multi-host inventory
  • configurable dependency discovery
  • notifications when a diagnosed incident requires approval

The long-term idea is for HomelabOps to act as a trustworthy operations assistant: autonomous enough to investigate complex failures, but constrained enough that humans remain in control of consequential changes.

Built With

Share this project:

Updates

Submission history