What it does

On-call is brutal. A bare alert at 3 AM means five minutes of dashboards, logs, and runbook hunting before you even know what's wrong. OnCall Copilot inverts that: it investigates the alert before it pages you.

When an alert fires, the copilot:

  1. Pulls the real metric from Prometheus and the error lines from Loki.
  2. Searches the web for the error string with Tavily.
  3. Asks NVIDIA Nemotron 3 Ultra (on Nebius Token Factory) to reason about the root cause against live evidence and your runbook.
  4. Pages you on Telegram with the likely cause, the evidence, and a suggested first action.
  5. Writes an incident report to persistent memory (markdown), so the next incident inherits the history.

Measured end to end: 74 seconds from injected failure to a page on your phone.

A real analysis it produced, unedited:

The demo-app is returning a 53.4% error rate (error_ratio_5m=0.534) with all recent errors showing "upstream dial timeout" (502), indicating the upstream dependency is unreachable or timing out. The low heap allocation (877,968 bytes) and inactive chaos flags (latency_mode=0) rule out local resource exhaustion or injected faults.

Using the heap and latency-mode evidence to rule out competing hypotheses, rather than restating the alert, is the behaviour that makes this worth building.

Why it's Personal AI

  • Always-on: a cron automation watches the alert stream every 30s.
  • Private & sandboxed: runs inside an NVIDIA OpenShell sandbox (NemoClaw) with a deny-by-default network policy. Only Token Factory, Telegram, Tavily, and your own Prometheus/Loki are reachable. Your telemetry never leaves your infrastructure except through the inference route you control.
  • Persistent memory: runbooks and incident reports live in plain markdown that the agent reads and updates, so each incident inherits the history.
  • Reusable skills: promql-query and logql-query work against any Prometheus/Loki.

How Nebius + NVIDIA were used

  • Nebius Token Factory — serves the reasoning model over an OpenAI-compatible endpoint. It dropped straight into NemoClaw's custom provider with four environment variables and no model-serving work at all.
  • NVIDIA Nemotron 3 Ultra (open model) — every root-cause analysis, reasoning over live metrics, log lines, a runbook excerpt, and web results.
  • NVIDIA NemoClaw + OpenShell — the sandboxed, policy-controlled agent runtime: deny-by-default egress, the managed inference.local route, cron automations, and the Telegram channel.
  • Tavily — web search for unfamiliar error strings during investigation.
  • Prometheus + Alertmanager + Loki + Grafana + Alloy on minikube as the demo stack.

Testing

Three layers, 38 tests. Go unit and race tests cover the chaos handlers and the relay queue and auth. An integration suite runs the real relay binary against stub Prometheus/Loki/inference, covering the failure modes that matter: token drift, an unreachable relay, batched alerts, path traversal in an alert label, and an unwritable lock. The end-to-end chaos harness (tests/e2e/run.sh) injects error, latency, and memory failures against the live stack and asserts a root-cause report for each — 3/3 scenarios pass.

Every fix was verified by reverting it and confirming the test fails. That caught a case where the copilot paged on schedule, on time, with a brief built on no evidence at all, while every unit and integration test still passed.

What's next

The incident memory is the interesting part. Right now each report is written and read back as markdown; the natural next step is for the copilot to cite prior incidents by similarity, so the third time a dependency times out it says so instead of reasoning from scratch.

Built With

  • alertmanager
  • docker
  • golang
  • grafana
  • grafana-alloy
  • kubernetes
  • loki
  • minikube
  • nebius-token-factory
  • nvidia-nemoclaw
  • nvidia-nemotron-3-ultra
  • nvidia-openshell
  • openclaw
  • prometheus
  • tavily
  • telegram
Share this project:

Updates

Submission history