Zero-Trust Edge Enclave is a privacy-first delivery workflow for confidential files. NVIDIA Nemotron 3 Super 120B on Nebius Token Factory, or Nemotron 3 Nano 4B on the local machine, advises on routing and follow-ups from five anonymous fields. The document, real identities and keys never reach the model, and fixed code validates every answer before anything acts.

Inspiration

The usual way to put AI into an internal workflow is to hand it the data and trust the prompt. In a regulated setting that option does not exist, and "the model was told not to look" is not a control.

So I narrowed the question. Instead of asking how much a model can do with full access, I asked the opposite: how much can it contribute when it is given almost nothing? Five anonymous fields. No document. No names. Not even the number of recipients.

What it does

Zero-Trust Edge Enclave delivers an encrypted document to an approved list of people, reports back which of them actually received it, and chases the ones who did not.

  • The sender's browser encrypts the file with AES-GCM before anything leaves it.
  • The recipient list is frozen into an immutable snapshot and confirmed twice.
  • Recipients are addressed by group codes regenerated on every revision, so a code never becomes a durable pseudonym.
  • Fixed code re-verifies identity, authorization, revocation, snapshot version, channel allowlist, expiry and retry budget before dispatch, and resolves recipients from the snapshot rather than from the model's answer.
  • Keys are wrapped per document and released against one-use tickets.
  • The sender picks a department and can untick anyone in it; only the ticked people are frozen into the snapshot.
  • Identity and access are shown apart: a person can be verified as signed in and still be refused a document they were not approved for, and the refusal is logged against the task without naming them.
  • The sender can open an evidence chain for each delivery: what was approved, the private mapping, exactly what each model call was given and answered, and whom fixed code actually reached.
  • Per-recipient receipts tell the sender which of the N approved people downloaded, verified and acknowledged.

Where the model actually decides something

Two decisions run through the same contract, each with its own five-field projection and its own validator.

Routing is the easy one, and I will not overstate it: given a prepared job and an approved channel set, a deterministic rule reaches the same answer. It is in the system because it is the simplest thing to check the boundary with.

Follow-up is the one that has no deterministic equivalent. A delivery that must be acknowledged has gone out, nobody has collected it, and the deadline is still days away. Wait, remind again, or raise it to a person? That depends on how time remaining, reminders already ignored, and partial collection sit against each other, and the right answer differs between tasks.

For that decision the model is told three things, none of which it can convert back into anything real:

  • timeCode — a position in the task's own window, so an identical value means a different hour on a two-day task and a two-month one
  • nudgeCount — 0, 1 or 2
  • pickupCode — PICKUP_NONE, PICKUP_SOME or PICKUP_ALL, never a count

It answers WAIT, REMIND or ESCALATE. It is never told who the recipients are or how many there are, and it does not decide who a reminder reaches — fixed code resolves that from receipts the model never sees, and silently passes over anyone who already collected.

Every answer is validated by fixed code before anything acts on it. An answer that fails validation pauses the delivery; an adviser that cannot be reached decided nothing, so the delivery waits and asks again, three times, before it pauses for a person. Once everyone has collected, the follow-up adviser is not asked at all.

How I built it

Two NVIDIA open models sit behind one contract. nvidia/nemotron-3-super-120b-a12b is served by Nebius Token Factory; nvidia-nemotron-3-nano-4b runs on the backend host through an OpenAI-compatible local runtime. The projections, the system boundaries, the output schemas and the validators are identical for both. Only the endpoint rule and the answer budget differ, so swapping a hosted 120B model for a local 4B one is a configuration change. Each model call is recorded with the projection it was given, the validated answer or refusal, and which outlet answered, so the sender can check the boundary rather than take it on trust.

Everything else is plain Node with no third-party runtime packages: one process serving four pages and the API from one origin, JSON-backed stores, an append-only audit trail, and a local key vault that refuses a directory readable by anyone else.

The judging instance runs the same code on Zeabur and calls Token Factory at runtime, behind a sign-in and under a USD 20 spending cap enforced in code, because the platform offers no per-key limit.

Challenges I ran into

An empty response that was not empty. The local outlet failed two of six cases with what looked like a transport fault. Diagnostics showed HTTP 200, finish_reason: "length", and zero characters of content. The model had spent its entire answer budget reasoning: 511 of 512 tokens, with 1940 characters of partial deliberation sitting in a separate field. Raising the budget fixed those two — and then a second, different problem appeared. Under a JSON schema constraint the model answers correctly in 55 tokens and seven seconds, but LM Studio assembles that answer into message.reasoning_content and leaves message.content empty. My client read only content, so it recorded a failure against a request that had succeeded. Reading both fields and constraining the output took the local outlet from 5 of 6 at 8.4–35.7 seconds to 6 of 6 at 3.7–4.1 seconds, with completion tokens falling from 645–1089 to 64–73.

The runtime changed under me. The day before recording, the local outlet went from about four seconds to between seven and twenty-seven. My request and prompt had not changed; LM Studio had started applying the JSON schema only after the model's reasoning, so the constraint no longer suppressed it. Turning reasoning off per request (reasoning_effort: "none") brought it to 2.3–2.7 seconds, but greedy decoding turned out to matter too: without temperature: 0, three of eighteen answers picked a reason the lookup table rules out, and the validator refused each.

A prompt written as prose. The follow-up prompt explained each field's ordering and each reason code's precondition in two paragraphs of English. Measured over 60 hosted calls, the model chose the right action nearly always but gave a reason contradicting its own input 11.7% of the time — reporting "no pickup yet" against a fully collected delivery. Rewriting the same facts as two lookup tables, changing no value, took reason coherence from 88.3% to 100% and p95 latency from 1728 ms to 1205 ms.

Numbers that were not the same kind of number. An early version labelled the window position TIME_1 through TIME_4. Next to a numeric reminder count, the model read the two as one scale and proposed a reminder at the very start of the window. Replacing the digits with words that count down what is left fixed it. I then tried the same on the reminder count and it made things worse — because that gain had come from two numbers competing, and with the window in words there was nothing left to compete with. A control that replaced the values with meaningless symbols, holding the table form constant, scored 76.7% against 86.7–88.3% for every labelled form: what a model needs here is a value it already understands, not a distinct one.

Retry scheduling had a bug my own tests caught. Full jitter can draw zero, which turns a backoff into an immediate retry. A separate issue scheduled retries after the download window had already closed; that case is now recorded as DELIVERY_WINDOW_CLOSED rather than mislabelled as exhausted retries.

Accomplishments that I'm proud of

The boundary holds under adversarial testing rather than by assertion. A fuzz run of 40,000 projections and 2,400 worker ticks against advisers returning hostile and malformed advice found zero invariant violations: no identity, group code, count or clock value ever reached a projection, no accepted advice left the allowlists, the reminder budget was never exceeded, and a fully collected delivery was never chased.

115 tests pass, thirty consecutive runs with no variation. Three local security scanners run against the exact published tree.

What I learned

The small local model was slower than the hosted large one. With reasoning turned off, the 4B local outlet takes 2.3–2.7 seconds where the hosted 120B takes 1.1–1.8. That inverts the usual assumption behind running the small one locally. And a local runtime can change underneath you: measure it again before you rely on an old number.

Where a model is not needed, say so. Routing would work as an if. Saying that plainly cost nothing and made it obvious which decision actually earns its place.

Constraining a model is a design position, not a limitation. Once the adviser could only see five fields and could only answer from a fixed set, the interesting work moved to what happens around it — re-verification, abstention handling, receipts, and who a reminder is allowed to reach.

What's next

Real transport instead of a dry run. Independent key custody, which is explicitly outside the current protection boundary rather than quietly assumed. And running the local outlet on hardware built for local inference, which the dual-outlet contract already makes a one-variable change.

Built With

Share this project:

Updates

Submission history