Inspiration

Fixing a broken service is slow, and most of the time goes to finding out why it broke, not to the fix. In a recent engineering-productivity assessment for a DevOps vendor, fault localization came up again and again as the biggest time sink. Engineers lose hours reading logs, events and dashboards before they write one line of code.

We wanted an agent that does the whole loop: diagnose, patch, verify, and hand over a PR a human can trust.

Mender grew out of Hermes, an earlier root-cause-analysis experiment of ours. Hermes supplied the idea, not the code: this repository is new and was built during the submission period.

What it does

Mender takes a broken Kubernetes service and returns a verified fix:

  1. Evidence. It collects events, pod status, logs and manifests with kubectl.
  2. Triage. A small model turns that raw evidence into a short list of signals and suspects.
  3. Root cause. A large model plans web searches, runs them through Tavily, and writes a root-cause report with numbered citations that a reviewer can open.
  4. Patch. A mid-size model writes the fix. It can change only the files it was given; a write to any other path is rejected.
  5. Verify. The tests run on the patched files in a Docker container with no network. If they fail, the model gets the test output and repairs the patch, up to three times.
  6. Hand-off. Mender prepares a pull request with the root cause, the evidence, the diff, the test results and the sources.

Nothing merges without a human. Mender's job is to make the review fast, not to skip it.

How we built it

  • Nemotron on Nebius Token Factory, split by task. All three models are called through one OpenAI-compatible endpoint. Measured over the 23-fault eval:

    • Nano 30B, triage: 25 calls; 280,173 input / 27,232 output tokens; 4.1 s median latency; $0.023.
    • Ultra 550B, search planning and root-cause report: 46 calls; 583,343 input / 80,196 output tokens; 4.6 s median latency; $0.824.
    • Super 120B, patch and repair: 30 calls; 367,524 input / 88,796 output tokens; 11.7 s median latency; $0.190.

The whole eval cost $1.04 in model calls. There is no fallback between tiers: a failed call is retried once and then raises an error.

  • Tavily is a real stage of the diagnosis, not an extra. The ultra model writes one to three queries, Tavily runs them at advanced depth, and the ranked results become the numbered citations in the report and in the PR. If Tavily fails after its retries, the diagnosis continues on cluster evidence and the report says so.
  • Sandbox. Each verification runs in a fresh Docker container with --network none, a memory limit and a CPU limit. We had planned to use Nebius Serverless Jobs for this and a Serverless Endpoint for the service. Token Factory does not have them, so the sandbox is local Docker behind a SandboxRunner interface, and the service is a plain HTTP server in a container image.
  • Eval harness. We injected 23 realistic faults into a kind demo cluster, in seven categories: app bugs, probes, configuration, secrets, resources, scheduling, and services and migrations. Each fault declares its expected labels, so each result is scored by code, not by hand.
  • Usage ledger. Every model call and every Tavily call is recorded with tokens, latency and cost for its tier. All numbers on this page come from that ledger.
  • Quality gates. CI runs ruff, mypy in strict mode, 84 tests, and a build of the service image that must start and answer its health check.
  • Open source under Apache-2.0. Setup is uv sync, two API keys in .env, then make demo. See the README.

Challenges we ran into

  • Runaway reasoning. The Nemotron models are reasoning models. With a small output budget, a call can end with finish_reason=length and an empty answer, because reasoning used the whole budget. In our first full eval, 4 of 23 cases failed this way. A larger budget, a different temperature and reasoning_effort did not fix it reliably. An instruction in the system message did, and we verified that fix against the exact prompt that failed.
  • Plausible but wrong fixes. A patch that makes the tests pass is not always a patch that fixes the fault. We score the two things separately: the fix passed in 20 of 23 cases, but the named root cause matched in only 9 of 23. Our label match is strict, so the second number understates the model, but the gap is real.
  • Sandbox safety and scope. A patch is rejected if it writes outside the allowed files or tries to escape through a path trick. This stopped two real out-of-scope patches in the eval. The PR step stages only the patched files, so unrelated work or credentials cannot enter a Mender branch.
  • No serverless compute on Token Factory. We found this during the build and had to change the design of the sandbox and the service.
  • No published price for one model. Nebius publishes no price table for each model, so the nano price in our config comes from third-party sites, and the config says so.
  • Our own demo was broken. The first demo video showed an error screen for its whole length, and the service image did not start. Both tools reported success. We now test the video in a real browser and start the image in CI.

Accomplishments that we're proud of

  • Mender's fix passed verification in 20 of 23 injected faults, and it named the correct root cause in 9 of 23 with a strict label match.
  • Median time from collected evidence to a prepared PR: 33.8 seconds.
  • The three-tier split cost $1.04 for the eval. The same tokens at the Ultra price would cost about $1.82, so the split saves about 43 %. This is a calculation from the measured tokens, not a second eval run.
  • We reproduced the runaway-reasoning failure on the live API, found its cause and verified the fix. Failures went from 5 to 3, and the 3 that remain have different causes.
  • Every number in the README, the demo video and this page comes from committed eval results. We removed the values that an earlier draft had invented.

What we learned

  • Verifying a fix matters more than generating it. The sandbox and the allowed-files check stopped 3 patches that would have looked like fixes.
  • A passing test is not a diagnosis. A reviewer still needs the root-cause report and its sources, which is why the PR carries both.
  • Small models are enough for reading. Nano handled triage of about 11,000 tokens for each incident at 6 % of the Ultra input price. The large model is worth its price only for the one reasoning step.
  • Reasoning models need room and a plan for silence. Give them a large output budget, and handle an empty answer as a normal case.
  • "Exit code 0" is not proof. Check the result itself: start the image, play the video, measure the diagram.

What's next for Mender: Kubernetes Incident-to-Fix Agent

  • More fault types and real incident traces beyond our injected set: multi-node networking, autoscaling and RBAC.
  • A better root-cause score that judges meaning, not only labels.
  • A remote sandbox runner, so verification does not need Docker on the same machine.
  • Integrations with alerting and on-call tools, so Mender starts from a page and not from a manual command.
  • Learning from merged and rejected PRs, so reviewer feedback improves later fixes.

Built With

  • kubernetes
  • nebius-ai-cloud
  • nebius-serverless-endpoints
  • nebius-serverless-jobs
  • nebius-token-factory
  • nemotron-3-ultra
  • nemotron-nano
  • nemotron-super
  • nvidia-nemotron
  • python
  • tavily
Share this project:

Updates

Submission history