Inspiration

Some software failures depend on timing. A normal test passes, but a delayed response or two overlapping actions can produce a different outcome. I wanted a way to show the difference with evidence, rather than ask an AI model to declare that a bug probably exists.

What PARALLAX does

PARALLAX starts two executions from the same application checkpoint. The control follows normal behavior; the counterfactual applies one bounded intervention. A deterministic verifier compares the resulting states against a security or correctness invariant.

The current release demonstrates two failures:

  • PAY-001: two actors interleave after reading the same cart. The control produces one payment; the counterfactual produces two payments for one ticket.
  • STOCK-001: two buyers read the last available ticket before either commits. The control sells one ticket; the counterfactual sells two and leaves negative stock.

Both experiments produce inspectable evidence and a PROVEN verdict.

How I built it

NVIDIA Nemotron 3 Super receives an application manifest and the same six executable intervention capabilities for both scenarios. It ranks three candidates and selects an experiment. Nebius Token Factory handles the model inference, while Nebius Sandboxes executes the selected intervention in isolated branches.

PARALLAX then checks the plan, execution, checkpoint, branch results and evidence hashes. The model proposes an experiment; it does not issue the final verdict.

The public Control Deck replays the captured executions. Its Verify Evidence button downloads the original plan, sandbox execution and assembled proof, recalculates their SHA-256 hashes in the browser, and checks the conditions behind PROVEN. It can also download a JSON verification receipt. Viewing the demo makes no paid API calls.

Challenges and what I learned

The hardest part was keeping a clear boundary between an AI hypothesis and a verified result. A plausible explanation is not a causal proof. I also had to make the selected model intervention match the action actually executed in the sandbox, and make the published evidence independently checkable.

I learned that a small set of well-defined, executable capabilities can be more useful than an unconstrained agent when the result needs to be reproducible and auditable.

What's next

I plan to test the approach against additional applications and connect counterfactual experiments to existing development and CI workflows. The current release focuses on two demonstrated failure families.

Built With

  • factory
  • fastapi
  • github-actions
  • nebius-sandboxes
  • nebius-token
  • nvidia-nemotron
  • python
  • react
  • vite
Share this project:

Updates

Submission history