Inspiration

An agent that writes a patch, runs its own tests and reports "done" is grading its own homework. Where the issue is ambiguous (how a tie rounds, what an empty list means), it picks one reading silently and its tests pass. More tests do not help until someone knows which requirement to ask about. Finding that requirement is the missing step; disagreement between independent attempts is a cheap way to find it.

What it does

  1. Three models generate candidate patches for the same issue through Nebius Token Factory.
  2. The candidates run on the same generated inputs in isolated ConTree sandbox branches. Every observation goes to an append-only trace; the runner never writes a verdict.
  3. A deterministic scorer reads only the trace. Candidates that fail a saved acceptance test are excluded (G4). If the remaining candidates disagree on a sampled input (G3), approval is withheld and the simplest separating input becomes a question. Model names and vote counts are hidden, and "not sure" is always an option.
  4. Before the user answers, one runtime Tavily search shows the relevant convention with its sources. It is displayed only; the verdict, the decision hash and the acceptance tests never read it.
  5. The answer is appended to a ledger, compiled into a typed acceptance test, and enforced in the next repair round.

With all gates enabled, PASS means the remaining candidate pool passed the regression checks and showed no disagreement on the valid sampled probes. It does not certify other inputs or establish correctness; candidates excluded by regression are not part of the approved pool. Everything not measured is marked in grey on the site.

How we built it

  • Token Factory (OpenAI-compatible endpoint). One chat completion per candidate slot with nvidia/Nemotron-3-Ultra-550b-a55b, Qwen/Qwen3-235B-A22B-Instruct-2507 and deepseek-ai/DeepSeek-V4-Pro, temperature 0.8, a per-slot seed. The reply is parsed into code, checked against an AST allow-list and hashed. Cross-model candidates were the diversity source that worked: one model sampled repeatedly detected the tie ambiguity in 1 of 5 runs, three models in 5 of 5.
  • NVIDIA Nemotron-3 Ultra fills the first candidate slot in every live session. In the 210-generation compliance experiment it followed a confirmed example placed in the prompt in 59 of 60 constrained generations; the two other models followed it in 26 and 21 of 60.
  • ConTree sandboxes (contree-sdk 0.3.6). Each candidate and each mutant gets its own branch node; probes run as disposable children. Sources are verified by remote SHA-256, every job id must be echoed back, and a sandbox failure is recorded as UNVERIFIABLE, never replaced by local execution. The whole clarification loop runs in sandboxes with loop.py --backend contree, as the demo session live8 did.
  • Tavily. research.py builds a deterministic query from the ambiguity axis, the issue text and the observed options (advanced search, 3 results, no LLM answer field) and records the result in the ledger between question and answer.
  • Scorer ported to JavaScript. site/scorer.js re-scores every published trace in the browser and compares its decision hash with Python's, so a reviewer can check the verdicts without trusting the build.

What we measured

Every number below comes from files in this repository. The published traces and experiment logs can be re-scored and re-aggregated with the commands in the README; the raw run directories that produced them are not committed, so a full rebuild from scratch is not reproducible here. Numbers marked recorded have no command.

  • Disagreement detection, 8 pure-Python tasks × 3 seeds, three models: 14 of 24 rounds returned NEEDS_CLARIFICATION (trace schema v2) and 15 of 24 (schema v1). The 48 cohort traces are published and re-score in the browser.
  • Prompt-constraint compliance, 7 conditions × 3 models × 10 generations, recomputed from the committed experiments/compliance_full.jsonl: each model's unconstrained rounding convention was fixed (Qwen 10/10 away-from-zero, DeepSeek 10/10 half-up, Nemotron 9/9 away-from-zero). With one confirmed example in the prompt the pooled satisfy rate was 13 to 19 of 30 depending on the target convention; with two examples, 13 to 24 of 30 (the same 210 generations, pooled by target convention instead of by model). Each cell holds ten generations, so differences between the lower rates are not resolved. Seven generations satisfied the examples while implementing a different policy.
  • G4 ablation on identical traces (python experiment_g4_ablation.py, fixtures only): in the constructed case A, where every candidate violates a confirmed decision the same way, G4 off gives PASS with three violators approved and G4 on gives CODE_INCOMPLETE. Case B shows no false block. How often case A occurs naturally was not measured.
  • Scorer fidelity: 96 of 96 published runs re-score to the Python hash in Node and in the browser (node site/test/scorer_test.mjs site/data). The build-time check of stripped versus full traces, 102 of 102, is recorded.
  • Live sessions on round_half: 8 LLM sessions with typed answers, 7 PASS (six in two rounds, one in three) and 1 UNVERIFIABLE where the regression gate left fewer than two distinct implementations, the fail-closed path working. The demo session live8 ran in ConTree sandboxes; three further sessions with scripted answers (ctree_llm1 to ctree_llm3) verified the sandbox path, 3 of 3 PASS with zero sandbox errors. All of these traces are published.
  • Cross-backend identity (4 groups, 9 runs, one decision hash per group) and the 49 canary checks are on the provenance page. Seven of the nine cross-backend traces are published and re-score in the browser; the canary results are recorded, a bounded leak check over the scanned outputs, not proof of complete isolation.

Challenges we ran into

The loop had never run in a sandbox. Single runs had executed in ConTree, but loop.py built the runner command without --backend, so every clarification session had silently run locally. An external review found it. The fix passes backend, workers and image through, records them in the session header, shows the execution environment per round on the session page, and refuses to re-run an existing session id. We then re-ran the loop in ConTree: three scripted sessions and the demo session live8. A later review found that --resume did not restore the recorded backend either; it now does, refuses a conflicting value, and a network-free regression test locks both behaviours.

Our first evaluation was circular. We had defined a "false pass" by the gate's own disagreement check, so the gate could only be measured against what it had itself produced. The redesign separates three programs that cannot read each other's data: the runner records observations only, the scorer never reads the reference tests, and the grader never reads the trace. Only then could the G4 ablation exist: the same trace, scored with the gate on and off.

What we learned

  • A fixed generation seed is not a reproducibility guarantee. Three sessions sent the same prompts with the same per-slot seeds on different days and each differed in one slot per round, while another round came back byte-identical.
  • One confirmed example does not identify a policy. Rounding -3.5 to -4 is consistent with three conventions.
  • The regression gate matters exactly where disagreement cannot help: when every candidate shares the same wrong reading.

What's next

Tasks beyond pure Python functions, more ambiguity axes through the live loop, a measurement of how often unanimous regression occurs in natural sessions, and an editor integration that asks the question where the developer already is.

Testing instructions

In the browser, no setup.

Locally, without any API key (Python 3.10; Node for the scorer check):

git clone https://github.com/minjun0208/behavioral-disagreement-gate && cd behavioral-disagreement-gate && pip install -r requirements.txt
python runner.py --task tasks/mean.json --run-id xb --gen hardcoded --backend local --probe-budget 20 && python scorer.py runs/xb/trace.jsonl cfg_full.json   # decision hash b03b4dd4…
node site/test/scorer_test.mjs site/data   # 96/96

With live models (what each key enables):

  • NEBIUS_API_KEY for Token Factory: --gen llm --models nvidia/Nemotron-3-Ultra-550b-a55b,Qwen/Qwen3-235B-A22B-Instruct-2507,deepseek-ai/DeepSeek-V4-Pro.
  • NEBIUS_PROJECT_ID with the key for ConTree: --backend contree (run python conformance_contree.py first; about 2.7 minutes per round on 8 workers).
  • TAVILY_API_KEY for the reference search; without it the search is recorded as unavailable and the loop continues.
  • A session id is used once: python loop.py --task tasks/round_half.json --session <new id> --gen llm --backend contree --models …

Feedback for Nebius

  • ConTree SDK 0.3.6: run() defaults to disposable=True, which discards the image a branch needs; node uuids come back as UUID objects rather than strings; a timeout is signalled as exit_code = -1 rather than a state; text uploaded from Windows needs newline="\n" or the remote SHA-256 will not match. conformance_contree.py in the repo runs eight checks before a session (auth and image resolution, disposable semantics, branch isolation, upload with SHA-256 verification, argv-only commands, result attributes and exit-code propagation, harness echo, timeout), so an SDK change is caught before it corrupts a trace. Documenting these behaviours would have saved a day.
  • Sandbox egress is open (EGRESS_OK measured from inside the sandbox). Isolation is filesystem and process isolation; a per-run network policy option would let gates like this one claim more.
  • Token Factory honours the seed parameter best-effort; the docs could say so explicitly.
  • Tavily's basic depth returned tutorials and off-topic pages for convention questions (relevance 0.17 to 0.24); advanced with three results was reliable. Domain boosting pulled in wiki revision diffs.

Built with

Python 3.10 · Nebius Token Factory (OpenAI-compatible API) · NVIDIA Nemotron-3 Ultra · Qwen3-235B · DeepSeek-V4-Pro · Nebius ConTree sandboxes (contree-sdk 0.3.6) · Tavily search API · vanilla JavaScript (ES modules, no build step) · GitHub Pages · ffmpeg and edge-tts for the video

Built With

Share this project:

Updates

Submission history