-
-
One change, two outcomes that disagree. Every test passes and the behaviour changed anyway. That contradiction is the whole problem.
-
What a reviewer actually sees: a small refactor, all checks green, merged. Nothing here is negligent and nothing here is a warning.
-
MUTINY runs the before and after versions on the same generated inputs. Where they disagree, the behaviour changed.
-
The GitHub Action comment on a real pull request, with the exact input that diverges and the before and after values it produced.
-
A merged refactor that stopped accepting an input the old code explicitly handled. Before returns a version, after raises a type error.
-
The same tool on a change that does describe itself. Behaviour changed, and MUTINY says so quietly. It comments only on surprises.
-
When nothing the generator wrote could build the object, it reads the class definition, builds it, and then does the job.
-
Install a checkpoint once, then fork it. 40 isolated microVMs in 3.7 seconds against 50.4 sequentially, for about half a cent a run.
-
The hosted version. Paste any public GitHub pull request URL and it fetches both versions, warms a sandbox and runs the comparison.
-
A live run on cachetools. A guard added to a constructor, behaviour changed, and the input that proves it. Three tenths of a cent.
-
MUTINY. Did that change preserve behaviour?
Inspiration
A pull request says "refactor, no behaviour change." The test suite passes. Two people approve it. It merges.
Three weeks later something is broken in production and the bisect lands on that diff. Nobody was careless. The tests passed because they were written against the old code, and the change did not break the paths they cover. The reviewers approved because the diff genuinely looks equivalent, and mostly it is. The AI review bot approved because it read the diff and formed an opinion. It never ran anything.
That is the actual pain point, and nothing in the toolchain addresses it. Every answer we currently have to "did this change behaviour?" is one of two things: a test somebody thought to write in advance, or a language model's opinion about code it never executed. Both fail in exactly the same place, on the input nobody imagined.
Here is how invisible it can get. shortuuid.int_to_string builds its output by concatenating characters and reversing the string at the end. A rewrite collected the characters in a list and reversed the list instead. For a single character alphabet those are identical. For any longer alphabet they are not. Every alphabet in shortuuid's own test suite is a string, so the suite is structurally incapable of telling the two versions apart. It passes. The behaviour changed anyway. No amount of reading that diff tells you this. Running it tells you immediately.
The obvious move is to run both versions on the same inputs and compare. Almost nobody does, because until recently it cost too much. You need both versions installed side by side. You need somewhere safe to execute a stranger's pull request. And you need inputs that actually reach the function that changed.
All three of those prices collapsed at once. Sandboxes made isolated execution fast and disposable. A language model made input generation nearly free. Something impractical for twenty years became something you can do in ninety seconds for half a cent.
So we decided to stop reading diffs and start running them.
What it does
Point MUTINY at a GitHub pull request. It downloads the code before and after, asks Nemotron for inputs that call the changed functions, runs every input against both versions inside a Nebius sandbox, and reports any input where the two disagree.
The output is not an opinion. It is an execution trace you can check in ten seconds:
Retrying(wait=None)._run_wait(state)
before 0.0
after TypeError: 'NoneType' object is not callable
That is from jd/tenacity#679, a real merged pull request titled "refactor: drop always-true truthiness checks". The maintainers merged it. The test suite was green. The author's reasoning is in the PR body and it is sound: wait is typed WaitBaseT, which implements neither __bool__ nor __len__, so the guard can only ever be true for a caller who respects the type. MUTINY found the exact boundary of that assumption.
It runs as a GitHub Action and says nothing on most pull requests. Of 49 merged PRs we measured, twelve changed behaviour and ten of those announced it in their own title. Reporting those is noise wearing the costume of diligence. A comment appears only when a behaviour difference is one the change does not mention.
Two rules, both about trust rather than capability:
- It never fails a build. A behaviour difference is information a reviewer weighs, not a verdict. A bot that blocks merges on its own judgement gets switched off within a week.
- One comment per pull request, edited in place. Pushing a fix replaces the warning instead of leaving it standing above a correction nobody scrolls to.
When a finding is real, one command turns it into a regression test the repository keeps, because otherwise the observation evaporates the moment someone merges.
And when a change removes behaviour the project's own documentation promises, Tavily finds the page and the comment quotes it. "Behaviour changed" in an undocumented helper is a curiosity. The same difference in something the docs specify is a broken promise.
How we built it
The one idea everything rests on: the model writes inputs, never assertions.
Ask a model for a test and it must produce an assertion, a claim about what the right answer is. Get that wrong and you have manufactured a false finding out of nothing.
Ask for an input and it cannot lie. A bad input is rejected identically by both versions and contributes nothing. A good one is executed twice and the results are compared by a machine with no opinion about which is correct.
That inversion is why this works. An earlier version that asked Nemotron to produce proofs scored 13% at both 120B and 550B. No amount of model size fixed it, because the task was wrong.
- Nemotron 3 Super 120B generates probes, rewrites functions, and writes the plain English explanation. Chosen by task shape, not size: it produced 42 usable inputs in 10.7s with zero reasoning tokens, where Nano spent 26,300 reasoning characters to produce 44, and Lightning produced 1.
- Nebius Sandboxes make it affordable. A checkpoint is installed once and then forked: 40 forks run in 3.7s against 50.4s sequentially, as microVMs rather than threads, so a probe that segfaults or deletes a file takes nothing with it. Executing a stranger's pull request is the whole point, and it never touches the runner.
- Tavily answers the question execution cannot: was the old behaviour promised to anyone?
When it cannot build the object, it goes and works out how. Some receivers need a paragraph of setup. TableDataElement wants a table, Job wants a scheduler, and asking for construction and behaviour inside one expression fails at the first, so the second never happens. So they are solved separately. Before any probe is written, a loop looks for a recipe, a snippet that builds the receiver and is proven by running it. Each attempt executes its candidates in the sandbox, reads what actually failed, and escalates the evidence it asks for: the class definition first, the repository's own tests only if that was not enough, because context costs tokens and should be fetched when the last attempt proved it was needed.
rich#4079 and schedule#604 went from nothing to twelve probes and a verdict, each built on the first attempt from the class definition alone.
It runs only when generation has produced nothing. Where generation already worked, the recipe makes things worse. cachetools.Cache.__setitem__ drops from 18 probes to 7, because anchoring every expression to one proven construction narrows what the model explores. An escalation that fires when nothing was wrong is a regression.
Challenges we ran into
Nondeterminism is the adversary, and it is relentless. Six separate causes of confident, specific, completely false findings, each one found by chasing a finding that turned out to be a lie:
| Cause | What it looked like |
|---|---|
| Memory addresses | every object without __repr__ "changed" |
| Set and dict order | identical frozensets, different repr |
| Hash seed | a real looking cache eviction bug |
random |
RRCache evicts differently each run |
os.urandom |
uuid4 differs by construction |
| Our own overlay | new code run against its old helper |
The last one is the one we are least comfortable with and most glad we found. A pull request changed two files. We overlaid only one, so the "after" side ran new code against an old helper, a state that exists in no commit, on no branch, in nobody's checkout. It reported a TypeError that no released version of that library can raise. The witness was real and the world it came from was not, which is worse than a wrong answer, because the evidence looks exactly as solid as a true finding.
We also publicly retracted a finding we had called "the one this project exists for." It was hash seed randomisation. That retraction is in the repository.
And a lesson about trust: the first live run of the GitHub Action reported "nothing to report, staying quiet" and went green, with empty credentials, having checked nothing at all. A green tick and no comment is exactly what a clean pull request looks like. A tool that reports total failure the same way it reports success is worse than no tool, because it is trusted and silent.
Accomplishments that we're proud of
- A real finding in a real merged pull request, reproducible by anyone in ten seconds.
- Zero false alarms across 14 agent refactors that preserved behaviour.
- A finding a full test suite could not see. The
shortuuidcase above: its own suite is structurally incapable of distinguishing the two versions. - It works on real frameworks. Django, 2,932 files, installs and runs end to end in 83 seconds for about half a cent.
- 40 of 45 pull requests with behaviour to check produce a verdict, at about half a cent each.
- Six documented classes of false finding, each with a reproduction, and one public retraction. We think that record is worth more than a longer feature list.
What we learned
Stability across runs is not the same property as being a value. Two fresh processes allocate identically, so a memory address looked perfectly stable within each version and different between them. No amount of re-running could catch it. Only canonicalisation could.
Asking a coin twice whether it is a constant gets "yes" half the time. shortuuid.random(length=1) draws one character from 57, took its one in 57 chance, and was reported as a behaviour change.
A model that contradicts itself is telling you something. Nemotron's own explanation read "the only observable difference is the object's address", directly beneath a verdict claiming behaviour had changed. The contradiction was the signal.
Measure the thing you are actually choosing between. Every model decision here came from a measurement, and the measurements repeatedly disagreed with the intuition. Bigger lost to smaller. More context lost to less. A helpful looking addition cost yield in the case that did not need it, twice.
What's next for MUTINY
Of 49 pull requests, 45 had behaviour to check and 40 produce a verdict. The five that do not are a decomposition problem. sqlalchemy's ClauseAdapter needs a selectable, which needs a table, which needs metadata. The construction loop escalates the evidence it asks for, but it still asks for the whole object at once instead of building the pieces in order. That is the next thing.
Beyond that: languages other than Python, and a confidence signal strong enough that a team could gate a merge on it.
Built With
- agentic-ai
- ast
- differential-testing
- fastapi
- github-actions
- llm
- microvm
- nebius
- nebius-sandboxes
- nebius-token-factory
- nemotron-3-super-120b
- nvidia-nemotron
- property-based-testing
- pytest
- python
- tavily
- upstash-redis
- vercel
Log in or sign up for Devpost to join the conversation.