Inspiration

Fixing a failing test is mostly trial and error: read the trace, guess a root cause, edit, run the tests, repeat. LLM assistants shorten the loop, but they still hand you a single suggestion - and a wrong guess costs another round trip. We wanted a tool that races several plausible fixes and keeps only the one its tests confirm.

What it does

FixFork takes a Python repo and a failing test command. It confirms the failure, runs one real Tavily web search for the failure signature, asks NVIDIA Nemotron 3 Super 120B (via Nebius Token Factory) for three divergent fix hypotheses with concrete edit plans, runs each candidate on its own forked branch, iterates the failing branches with Nemotron 3 Nano 30B (up to two rounds per branch), and keeps the branch whose tests pass - ties go to the smallest change. Losers are rolled back, never merged. If no branch goes green, no patch is exported; the report names the least-failed branch as a lead, not a fix. Output: a git apply-able patch, a self-contained HTML race report (with per-branch token and USD accounting), raw model replies, and a JSONL event log.

How we built it

  • Python 3.10+, standard library only - our environment blocks package installs, so everything is plain REST against Token Factory's OpenAI-compatible API.
  • A model router with reasoning-budget retry: when a reasoning model exhausts max_tokens, content comes back empty with finish_reason=length; the router retries with a doubled budget (4096 -> 8192 -> 16384 -> 32768).
  • A git-style sandbox abstraction (checkpoint / fork / rollback) with two backends: local temp-dir branches (default, no isolation) and the live Token Factory Sandboxes backend (opt-in), where fork/rollback move content-addressed state images - a full three-branch race ran in the sandboxes on 2026-09-29.
  • Evidence-based judging: passing tests first, then fewest lines changed; per-branch cost tracking.

Challenges we ran into

  • Reasoning-budget exhaustion is silent unless you know the pattern - we found the retry by experimenting.
  • Sandbox access states are opaque from outside: with a Project header present, a wrong id, a missing role and missing beta access all look identical (HTTP 200 with every permission false, or 403 Insufficient permissions: list).
  • Package installs were blocked in our environment, so the HTTP client, patch export and test-runner glue are all standard library.

Accomplishments we're proud of

  • A recorded demo run turned 2 failing tests into a passing 1-line fix in 26.4 s end-to-end for about $0.003 in model API cost, and the exported patch re-applied cleanly with git apply on a pristine checkout.
  • 156 unit tests ship with the repo; the pipeline was validated through real end-to-end runs during development, including a later run with live Tavily web grounding (five sources).
  • The test suite is the referee and is protected: proposed edits to test files, CI workflows or build/config files are refused before any sandbox work (blocked), so such a branch can never win. A regression test replays a recorded race in which a test-editing branch used to score green - it is now refused while the real fix still wins; the exported patch is re-checked so renamed or quoted protected paths are withheld (test-path conventions are matched case-insensitively: Tests/, __tests__/, testdata/).
  • A full three-branch race ran entirely inside Nebius Token Factory Sandboxes (2026-09-29): the baseline failed inside the VM, three branches forked from one state, and the winning 1-line patch was exported from the sandbox VM and re-applied cleanly outside with git apply; 13 sandbox operations measured $0.0044 in sandbox cost.
  • A run against a real external repository (python-humanize, 2026-09-29) turned 6 failing tests into a winning patch; re-applied with git apply outside the pipeline, the repository's own test suite then reported 310/310 passing (dev-only dependency files excluded).
  • A real-world result (merged 2026-09-30): pointed at a live upstream issue in python-poetry/tomlkit (#619, stdlib fold support), FixFork's runs reproduced the failure and located the fix area; run traces are public (examples/tomlkit-619/ in the repo). Neither run finished the patch on its own; the final patch - 10 source lines plus 2 regression tests, completed and verified by hand from the run report - passed tomlkit's own full suite (1,060 passed) and was merged into tomlkit master by the maintainer 66 minutes after the pull request opened (PR #620: +52/-1 across 2 files, 24/24 CI checks green).
  • One Tavily search per run is a real runtime call - and the tests remain the only verdict.

What we learned

  • Evidence beats eloquence: letting tests pick the winner lets you accept fixes you would not trust from prose alone.
  • Cost transparency changes behavior: per-branch token accounting makes "is this still worth it?" an explicit decision.
  • Good retries matter more than good prompts when models think silently.

What's next

  • Token Factory Sandboxes backend for real branch isolation on Nebius infrastructure - now wired and verified live; next it becomes the default backend.
  • More languages beyond Python; packaging as a GitHub Action.

Requirements

  • Python 3.10+ and git on PATH; standard library only - no pip install needed.
  • Live runs need internet access and a Nebius Token Factory API key. The Tavily search works keyless by default; an API key is optional.
  • The repository README has full setup and quickstart instructions, including an offline demo mode (deterministic local router) that needs no API keys at all.

Links

Built With

  • ai-agents
  • developer-tools
  • git
  • llm
  • nebius-token-factory
  • nvidia-nemotron
  • python
  • tavily
  • unit
Share this project:

Updates

Submission history