Recovery Agent
Recovery Agent is a control-host recovery runtime for Linux services. It gives developers and operators bounded AI-assisted recovery without handing a model anonymous shell access to production.
Small teams cannot watch every service manually, but "give the agent a terminal and hope" is not a safety model. Recovery Agent separates deterministic infrastructure authority from model reasoning.
How it works
The control host owns the recovery workflow. Production nodes expose only narrow typed health and recovery operations.
Deterministic control owns configured targets, health probes, restart budgets, write-ahead durability, outbound mTLS transport, systemd mutation, fresh verification, and audit.
Strands reasoning handles bounded triage, specialist analysis, synthesis, proposal planning, veto-style critique, semantic watch compilation, and postmortems.
The model never receives a generic recovery shell and never owns human approval.
What it does
Recovery Agent provides:
- fleet, node, and service inspection;
- deterministic health sweeps and application probes;
- bounded typed restart_service recovery;
- Recovery Readiness classification;
- service, node, certificate, and deployment watches;
- natural-language semantic watches compiled once into deterministic recurring schedules;
- dependency-aware rolling restart budgets;
- durable incidents, plans, watch definitions, deployment causality, and append-only audit history;
- outbound TLS 1.3 mutual-authenticated node transport;
- named local operator approval outside model-facing MCP;
- bounded Strands triage, specialist, synthesis, planner, critic, and postmortem roles.
Recovery is evidence-driven
A mutation is treated as a transaction with evidence:
unhealthy target -> durable intent checkpoint -> one typed recovery -> fresh verification -> durable outcome
If the intent checkpoint fails, zero recovery mutation occurs. Ambiguous execution after a crash is not blindly replayed. Dependencies can block downstream recovery without consuming restart budget, and every successful mutation is followed by fresh deterministic verification.
Production boundary
Production nodes initiate outbound TLS 1.3 mutual-authenticated sessions to the control host. Remote/model callers cannot supply unit names, shell commands, health-check URLs, TLS targets, arbitrary ports, credentials, or deployment-marker paths. Those remain configured product state.
Human approval uses named principals over an owner-only local Unix socket and stays outside model-facing MCP.
Why it matters
Operations failures create exactly the kind of work agents should help with: repetitive inspection, correlation, diagnosis, and recovery planning under time pressure. Recovery Agent lets the AI reason where reasoning is useful while keeping infrastructure authority narrow, typed, durable, and auditable.
Strands
Strands performs bounded triage and specialist reasoning after deterministic recovery logic has gathered evidence. The planner can return only restart_service or none, and a veto-only critic may reject the proposal but cannot invent a new action. Model output cannot select another target, approve a plan, or mutate a node directly.
Hackathon fit
Recovery Agent targets the Professional Agents track. It is designed for developers, operators, and small technical teams who need safer, faster recovery from routine service incidents without replacing operational controls with unrestricted AI authority.
The public repository includes source code, setup instructions, an MIT license, styled architecture and safety diagrams, a demo runbook, packaged acceptance paths, and validation evidence.
Visual architecture
System architecture
Recovery safety boundary
Built With
- codex
- java
- mcp
- node.js
- postgresql
- strands-agents
- systemd
- tavall-database
- tls-1.3
- typescript
Log in or sign up for Devpost to join the conversation.