Inspiration

Anyone who has been on call knows the first ten minutes of an incident are the worst. A page arrives: "no space left on device on web-server-01." You SSH in, run df -h, and it says 41% used. The obvious answer is wrong. Twenty minutes later you remember df -i and find the inode table at 99%, because something quietly wrote forty thousand tiny files into a mail queue. The diagnosis is the slow part, it happens under pressure, and it is usually done by whoever is least awake.

That is a good place for an agent and a terrible place for an autonomous one. Two things stopped us, and they are the reason this project exists in the shape it does. First, nobody sensible hands an LLM write access to production. Second, nobody wants to paste server credentials into a chat window, because that is how credentials end up in someone else's logs.

So we built the thing we would actually run: an agent that can only look, tells you what is wrong with evidence attached, and cannot change anything until a human says yes — and whose credentials never leave the machine they were typed into.

What it does

RootCause AI takes an unstructured operational alert, runs strictly read-only diagnostics on the affected host, pinpoints the root cause from real evidence, and produces a structured incident report. Anything that changes system state waits behind an approval prompt.

A real run against our testbed, which is a container with a /var/spool/postfix tmpfs capped at 50,000 inodes and ~40,000 tiny deferred-mail files:

  • It reads df -i at 89% IUse while df -h shows only 62% of bytes used — the alert said 41%, and the bytes were never the problem.
  • It runs a directory density scan and attributes 40,000 of those inodes to maildrop/queue1 through queue8, 5,000 files each.
  • It names the mechanism: a fixed inode pool, nearly depleted by a tiny-file workload, so new queue files fail with ENOSPC even though space looks free.
  • It proposes the minimum fix — clearing those queues — and stops. A red panel shows the exact command, the host, and the blast radius. It runs only on a y, and a n is recorded in the report instead.
  • It writes up the incident as a structured report: severity, symptoms, root cause, quoted evidence, impact, ordered remediation with per-step risk and an approval flag, and preventive next steps.

It runs over Docker or SSH, as an interactive terminal UI (rc), as a one-shot command for cron and CI, or as a Python API. Around the diagnostics sit six handlers that mediate every model and tool call: a secret scrubber, guardrails, a rate limiter, skill steering, a tone guard, and the human-in-the-loop gate.

How we built it

Built on the Strands Agents SDK in Python, with Amazon Bedrock as the default model backend and a provider layer that also supports the Anthropic API, the Command Code gateway, and any OpenAI-compatible endpoint, including a local one.

The security model is the interesting part. The model never receives a credential: tools address hosts by alias, and targets.py resolves that alias to a user, host and port locally. Five independent layers keep it that way — credentials are parsed out of typed text before a message exists, redacted from every message on every model call, redacted from command output before it enters the conversation, redacted inside the session manager's write path so nothing lands on disk, and masked in everything rendered to the screen. Passwords are never stored: SSH keys are tried first (ssh-agent, ~/.ssh/*, ~/.ssh/config aliases), a password is requested only if that fails, and it lives in memory for that session only. Host keys are verified strictly, with an explicit fingerprint confirmation on first contact.

We leaned on the SDK's own primitives rather than inventing policy: InterventionHandler with its Deny and Guide control actions, and the vended HumanInTheLoop handler for the approval gate. Read-only enforcement sits in the transport, where every command is checked against a whitelist and written to a redacted audit log; remediation is the single code path that lifts that restriction, and it is only reachable after approval.

The code is organised by responsibility — cli/, agent/, transport/, security/, diagnostics/, providers/, skills/, reporting/, ui/ — so the guarantees live in security/ and the host access lives in transport/, rather than being spread across a flat pile of modules.

It ships as one self-contained executable. PyInstaller bundles CPython and every dependency, nfpm produces .deb and .rpm packages that declare no dependencies at all, and the binary is built inside Ubuntu 22.04 so the glibc floor admits newer Ubuntu, Fedora and RHEL. A bare container with no Python and no ca-certificates can run it, because TLS trust comes from a bundled certificate store.

Testing was treated as the product, not the afterthought: 201 unit tests plus seven end-to-end tests that drive the real agent loop — same handlers, same tools, same session manager — against a real container, using a scripted model so no API key is needed. Those tests assert against container state and captured command output rather than model prose: after a denial the files are still there, after approval they are gone, and a mkfs proposal is refused even with y already queued.

Challenges we ran into

APIs that fail silently. The SDK accepts a model config, warns about unrecognised keys, and drops them. Our temperature=0.0 — the thing that makes triage reproducible — was being discarded on two of four provider paths. We only caught it by constructing the real SDK classes in tests and asserting the value landed in the config.

Hook ordering is architecture. The session manager registers its hook before any user hook, so a redaction hook added alongside it would have written the secret to disk first and redacted afterwards. The fix was to redact inside the session manager's write path, not beside it.

A module name that broke the packaged binary. Naming our credential module secrets.py shadowed the standard library module that paramiko imports for cryptography. Unit tests passed — pytest puts the repository first on the path — while the installed binary hit the real stdlib and crashed on startup. Renamed, and now the tests construct the shipped artifact.

Self-containment is only provable in a bare environment. The binary worked everywhere until we ran it in a minimal image, where it failed with APIConnectionError. The system certificate store did not exist, and httpx reported a TLS failure as a connection error. We now point TLS at the bundled store explicitly.

The agent had to stop guessing. In one run the transport itself was broken, and the honest answer was "I cannot see anything and I will not speculate." Getting that behaviour required the system prompt, the tools and the instrumentation to agree that unevidenced claims are worse than an admission of blindness.

Domain vocabulary defeats naive heuristics. Our tone guard flagged ENOSPC, SIGTERM and OPERATOR DENIED as shouting, forcing a pointless rewrite after the report had already been produced. And the skill classifier picked "service failure" over "inode exhaustion" because generic words like failing and stalled outscored the precise symptom no space left on device. Both now weight distinctive signals over common ones.

An approval gate that ate the input. With piped stdin, the approval prompt consumed the next line of input — a follow-up question — as its answer, then looped until the run died. A safety gate that fails in the wrong direction is worse than no gate, so it now denies outright when there is no interactive terminal.

Accomplishments that we're proud of

The approval gate is real, and we can prove it. The end-to-end tests check the container's filesystem, not the agent's description of it: deny and the 5,000 files are still there; approve and they are gone; propose mkfs.ext4 and it is refused before a human is even asked.

Credentials demonstrably never leave the machine. One test asserts the password is absent from the exact messages handed to the model; another walks every file in the session store; two more verify that a secret arriving via command output is masked before it enters the conversation. There are no secrets in the registry file at all.

Triage is deterministic. Greedy decoding, fixed diagnostic playbooks injected per failure family, and per-turn budgets that force convergence — the same alert and the same evidence produce the same path, which is what you want from a first responder.

It diagnoses like a sysadmin, not a search engine. The moment that sold us on the design was the agent explaining why df -h looked healthy: it had read the inode table, counted files per directory, and found the one directory tree responsible.

One binary, no dependencies. Verified by installing the packages and running the binary on Ubuntu 22.04, Ubuntu 24.04 and Fedora — images with no Python, no pip and no CA store — and it starts, self-checks and runs.

What we learned

  • A guarantee you have not tested negatively is a hope. "We redact secrets" meant nothing until we grepped the disk and inspected the model payload in a test.
  • Silent failure is the hardest class of bug to design against. If a library warns and continues, your tests must assert on the object the library actually uses.
  • The agent's refusal to fabricate was worth engineering for. Failure paths — no credentials, a broken transport, a tool that returns nothing — are where trust is won or lost.
  • Ordering, not features, drives security. Four of our five credential defences exist because of when something runs, not what it does.
  • Domain vocabulary breaks generic heuristics. A tone checker written for prose misreads an incident report; a classifier written for English misreads an alert.
  • Building for real Linux means building on real Linux: compiled wheels and glibc versions make "it works on my machine" a packaging problem, not a code problem.

What's next for RootCause AI

Webhook-triggered triage. The one-shot mode already runs without a terminal, so what is missing is the ingress. We want a small service that accepts alert payloads from the tools people actually run — Alertmanager, PagerDuty, Datadog, or a CloudWatch alarm routed through EventBridge — verifies the sender's HMAC signature, deduplicates on the alert fingerprint, and starts a triage run, posting the report back to whatever raised the alert. Alert text and log output are untrusted input, so that path also needs input screening in front of the model and a check that no instruction arriving in a log line can change what the agent is allowed to do; the tool-side validation we already have is the backstop, not the whole answer. The approval gate cannot block on stdin in a service, so this runs on the SDK's interrupt/resume: a pending remediation is durable state with an id, a human resolves it over HTTP or Slack, and the loop continues where it stopped. Deployed as a container behind an ALB, or as a Lambda for low alert volume.

Web UI. A browser client over the same agent: live transcript, each tool call with its arguments, and the approval request rendered inline with the exact command and its blast radius. Server-side this is the interrupt/resume path above rather than a thread parked on a console prompt, so an approval is a record — who approved what, and when — and more than one person on call can see it. The terminal interface stays; this is for teams that do not live in one.

TUI improvements. The interface works, but a few things still show its origins:

  • Approval deserves a real modal with a command preview, instead of an answer typed into the same input box used for questions.
  • Transcript paging and search: a long triage currently scrolls out of a fixed buffer, and the evidence is worth keeping on screen.
  • Switching targets, and opening a past session, without restarting the program.
  • Severity colours that hold up on a colour-blind palette, and a plain mode for terminals that strip ANSI.
  • A compact transcript mode — one line per tool call — for narrow terminals and for recordings.

None of this changes the constraint the project is built on: the agent reads, and a person decides.

Built With

Share this project:

Updates

posted an update —

The honest limitation until today: four fixed tools meant four diagnosable failure families. Ask about anything else and it correctly said "I cannot see that, and I will not guess" — accurate, but not useful.

run_readonly_command changes that: the agent can now run one read-only command of its own choosing, while the four fixed tools remain for the common cases. The whitelist grew from 32 commands to 87 for normal Ubuntu and RHEL hosts — ps, du, dmesg, journalctl, ss, netstat, ip, lsof, iostat, dpkg, rpm, awk, sed, jq. The blocklist grew to 142 tokens and now includes shells, interpreters and network clients, with new guards for write flags (sed -i, journalctl --vacuum-time, dpkg -i, ip addr add), read-only subcommand allowlists for systemctl and service, and raw-string checks for command substitution, pipes into a shell and file redirection.

Relaxing the surface required tightening the enforcement: the blocklist is now applied only at command positions, because a naive token scan cannot tell /etc/passwd (a file being read) from /bin/rm (a command being run), and the loop-level guardrail now enforces strictness per tool — only propose_remediation may carry a state-changing command.

The same run that proved it found three more bugs: our own density scan used $(...) and no longer passed the policy, the agent was not told which host the CLI had selected, and a turn that hit the output token limit crashed instead of continuing.

Same host, same model, memory-pressure alert, before and after:

Before: "There is no ps/top-style probe in my toolset… I cannot name the leaking process from free -m alone, and I will not guess."

After: ran ps aux --sort=-rss, cat /proc/meminfo, /proc/pressure/memory, the container's cgroup accounting and dmesg | grep -i oom; ruled the container out on cgroup evidence (memory.current 207 MB, oom_kill 0), confirmed host pressure independently (AnonPages 4.7 GB, Committed_AS 49 GB against a 20 GB commit limit), and reported that the consumer sits outside the visible PID namespace — with the exact command to run where it is visible.

Log in or sign up for Devpost to join the conversation.

Submission history