Inspiration

I asked my assistant a simple question: "What is the tower?"

It answered with the replication cycle of HIV.

Nothing was broken. The matching worked exactly as written — the French word tour was hiding inside détournement, in a medical document it had ingested. It found a match, and it answered with total confidence.

That is the failure I wanted to fix. Not "the model is not smart enough" — the model never says it doesn't know. Every assistant I use will invent a plausible answer rather than admit a gap. For a companion that sits on your desktop all day and reads your real machine, that is not a quirk. It is the whole problem.

So I built Arthur around one rule: he would rather say nothing than say something wrong.

What it does

Arthur is an always-on companion that floats on your desktop. He has a face, a chat bar, a voice, and a transparent window so your work shows through him. He answers in five layers, and he only climbs when the layer below has nothing honest to say:

Layer What it is Measured
0 Tools — he looks at the real machine 3–5 ms
1 Written rules — whole-word matching, no substrings 3 ms
2 Ingested documents ~40 ms
3 NVIDIA Nemotron on Nebius — the large model 891 ms
4 Admission — "I don't know, and here is why" 1 ms

The point is layer 4. It is not a fallback. It is a destination.

Ask him about health and he refuses outright: "I don't handle health questions — I was made for the tower, not for doctors." Ask him the colour of a neighbour's hat and he says he cannot know. Break his connection to the server and he says "I could not look" — he never fills the silence with a guess.

His tools read real state, never a cached claim:

  • qui_est_en_ligne — which containers are running on the VPS, which machines answer on the home network. Measured at call time.
  • serveur_va_bien — uptime, load, disk, memory, dead containers, and whether the public pages actually open. Then he decides: fine, or not fine and why.

The second tool found a real outage the first time I ran it: the tower's dashboard was returning 404. I did not know.

Every tool is also exposed to the browser through WebMCP, so any agent looking at the cockpit page can call them too. The list is not hardcoded — the page asks the machine what tools exist. Arthur gains a tool tomorrow, it appears by itself.

How I built it

The engine is a rules engine, not a model. Rarity-weighted whole-word matching over an indexed corpus, in pure Python. Procedures need at least two matching words to fire, and any word appearing in three or more procedures is dropped from the procedure index entirely — otherwise one common word drags a whole document into an unrelated answer.

Nebius is the fifth layer, not the first. Arthur calls nvidia/Nemotron-3-Ultra-550b-a55b on Nebius Token Factory only when his own layers have nothing true to offer. I benchmarked all four NVIDIA Nemotron models Nebius serves, on three probes — one easy, one arithmetic, one trap where the only correct answer is "I cannot know":

Model Correct Mean latency
Nemotron-3-Ultra-550b-a55b 3 / 3 891 ms
Nemotron-3_5-Lightning 2 / 3 728 ms

Ultra is 163 ms slower and does not guess. For this project that trade is not close.

The face is a native GTK3 window with an RGBA visual, because a browser renders transparency as white. The voice is Piper (ONNX), running on-device. No pixels leave the machine.

Challenges I ran into

1. Nebius returns an empty content on reasoning models — silently.

This one cost me an hour and is worth reporting. nemotron-3-super-120b-a12b spends its budget on reasoning first. Under a normal max_tokens, the reasoning fills the budget, the answer never gets written, and the response comes back with content: null and finish_reason: "length". Code using the standard OpenAI client crashes on NoneType has no attribute strip.

I measured it across all four models and three budgets:

Model 100 tok 300 tok 900 tok
Nemotron-3-Ultra-550b-a55b 205 chars 205 chars 205 chars
nemotron-3-super-120b-a12b EMPTY EMPTY 153 chars
NVIDIA-Nemotron-3-Nano-30B-A3B 151 chars 151 chars 151 chars
Nemotron-3_5-Lightning 399 chars 1244 chars 149 chars

Arthur now handles it explicitly: when content is empty he says "the large model thought so long it had no room left to answer" — and he never hands back reasoning_content as if it were an answer. Those are its scratch notes, not its claim.

2. Two wrong answers that looked like bugs but were design flaws.

"How much is 17 times 23?" returned a procedure about reminders. Cause: "how much" and "are" are stop words, leaving only "times" — which appears in a procedure. And Arthur had no calculator at all. He has one now, and it parses digits itself rather than evaluating the sentence.

"Explain in one sentence what a transistor is" returned a procedure about onboarding new users. Cause: "explain" and "sentence" both appear in procedures. Two matches were enough to win. But those words describe how to answer, not what about. I removed 33 of them from scoring.

Neither was found by reading the code. Both were found by tests.

3. The model I had been calling did not exist.

Nebius returned 404, which I first read as an auth problem. It was not — the model id had simply been retired. Listing /v1/models and reading what is actually served took thirty seconds and saved an hour of guessing at the wrong layer. A lesson Arthur is built around, applied to me.

Accomplishments I'm proud of

82 green checks, and half of them test refusal.

Suite Result What it proves
Calculator 25 / 25 7 of them are cases he must decline
Eyes on the server 27 / 27 includes breaking SSH on purpose
Tools exposed 19 / 19 unknown tool → 404 in plain French
Real browser 11 / 11 the JS actually executed, a tool actually called

The suite I care about most deliberately breaks the connection to the server and asserts three things: he admits he could not look, he invents no container name, and he does not say "everything is fine". A gate that has never refused anything is not guarding anything.

The browser suite runs the WebMCP file in a real JS runtime with a fake document and a live fetch, registers the tools, and calls one. Not "the file exists" — the tool answered.

What I learned

A wrong answer is worse than no answer, and it is much harder to detect. When Arthur told me about HIV, nothing logged an error. The matching worked. Only a human reading the output could tell.

My own tests lied to me twice today. One counted an off-topic answer as correct because the word "none" appeared in it. Another compared the length of a string to 60 instead of the string. Both were green. A green test you have not tried to break is a decoration.

Reasoning models break the OpenAI contract in a way that fails silently. content: null with finish_reason: "length" is not an error — it is a successful response that carries nothing. Every integration that assumes content is a string will crash on it.

What's next for Arthur

  • Mini Arthur for the public site — all the knowledge except anything private, no tools, no disk writes.
  • Acting, not just answering: fix a server fault, deploy infrastructure, build an app — reading studies directly from the server's files.
  • A plugin system so anyone can teach him their own tools.
  • The transporter: a USB key that deploys Arthur, with all his tools and guard-rails, onto any machine with one script.

Built With

Share this project:

Updates

Submission history