Inspiration
Fixing a failing test is mostly trial and error: read the trace, guess a root cause, edit, run the tests, repeat. LLM assistants shorten the loop, but they still hand you a single suggestion - and a wrong guess costs another round trip. We wanted a tool that races several plausible fixes and keeps only the one its tests confirm.
What it does
FixFork takes a Python repo and a failing test command. It confirms the failure, runs one real Tavily web search for the failure signature, asks NVIDIA Nemotron 3 Super 120B (via Nebius Token Factory) for three divergent fix hypotheses with concrete edit plans, runs each candidate on its own forked branch, iterates the failing branches with Nemotron 3 Nano 30B (up to two rounds per branch), and keeps the branch whose tests pass - ties go to the smallest change. Losers are rolled back, never merged. If no branch goes green, no patch is exported; the report names the least-failed branch as a lead, not a fix. Output: a git apply-able patch, a self-contained HTML race report (with per-branch token and USD accounting), raw model replies, and a JSONL event log.
How we built it
- Python 3.10+, standard library only - our environment blocks package installs, so everything is plain REST against Token Factory's OpenAI-compatible API.
- A model router with reasoning-budget retry: when a reasoning model exhausts
max_tokens, content comes back empty withfinish_reason=length; the router retries with a doubled budget (4096 -> 8192 -> 16384 -> 32768). - A git-style sandbox abstraction (checkpoint / fork / rollback) with two backends: local temp-dir branches (default, no isolation) and the live Token Factory Sandboxes backend (opt-in), where fork/rollback move content-addressed state images - a full three-branch race ran in the sandboxes on 2026-09-29.
- Evidence-based judging: passing tests first, then fewest lines changed; per-branch cost tracking.
Challenges we ran into
- Reasoning-budget exhaustion is silent unless you know the pattern - we found the retry by experimenting.
- Sandbox access states are opaque from outside: with a Project header present, a wrong id, a missing role and missing beta access all look identical (HTTP 200 with every permission false, or 403 Insufficient permissions: list).
- Package installs were blocked in our environment, so the HTTP client, patch export and test-runner glue are all standard library.
Accomplishments we're proud of
- A recorded demo run turned 2 failing tests into a passing 1-line fix in 26.4 s end-to-end for about $0.003 in model API cost, and the exported patch re-applied cleanly with
git applyon a pristine checkout. - 156 unit tests ship with the repo; the pipeline was validated through real end-to-end runs during development, including a later run with live Tavily web grounding (five sources).
- The test suite is the referee and is protected: proposed edits to test files, CI workflows or build/config files are refused before any sandbox work (
blocked), so such a branch can never win. A regression test replays a recorded race in which a test-editing branch used to score green - it is now refused while the real fix still wins; the exported patch is re-checked so renamed or quoted protected paths are withheld (test-path conventions are matched case-insensitively:Tests/,__tests__/,testdata/). - A full three-branch race ran entirely inside Nebius Token Factory Sandboxes (2026-09-29): the baseline failed inside the VM, three branches forked from one state, and the winning 1-line patch was exported from the sandbox VM and re-applied cleanly outside with
git apply; 13 sandbox operations measured $0.0044 in sandbox cost. - A run against a real external repository (python-humanize, 2026-09-29) turned 6 failing tests into a winning patch; re-applied with
git applyoutside the pipeline, the repository's own test suite then reported 310/310 passing (dev-only dependency files excluded). - A real-world result (merged 2026-09-30): pointed at a live upstream issue in python-poetry/tomlkit (#619, stdlib
foldsupport), FixFork's runs reproduced the failure and located the fix area; run traces are public (examples/tomlkit-619/ in the repo). Neither run finished the patch on its own; the final patch - 10 source lines plus 2 regression tests, completed and verified by hand from the run report - passed tomlkit's own full suite (1,060 passed) and was merged into tomlkit master by the maintainer 66 minutes after the pull request opened (PR #620: +52/-1 across 2 files, 24/24 CI checks green). - One Tavily search per run is a real runtime call - and the tests remain the only verdict.
What we learned
- Evidence beats eloquence: letting tests pick the winner lets you accept fixes you would not trust from prose alone.
- Cost transparency changes behavior: per-branch token accounting makes "is this still worth it?" an explicit decision.
- Good retries matter more than good prompts when models think silently.
What's next
- Token Factory Sandboxes backend for real branch isolation on Nebius infrastructure - now wired and verified live; next it becomes the default backend.
- More languages beyond Python; packaging as a GitHub Action.
Requirements
- Python 3.10+ and
giton PATH; standard library only - nopip installneeded. - Live runs need internet access and a Nebius Token Factory API key. The Tavily search works keyless by default; an API key is optional.
- The repository README has full setup and quickstart instructions, including an offline demo mode (deterministic local router) that needs no API keys at all.
Links
- Code: https://github.com/tuyentran4992/fixfork
- Narrated walkthrough: https://youtu.be/A6xS5TFVJs4
- Demo page: https://fixfork.netlify.app
Log in or sign up for Devpost to join the conversation.