Inspiration

A plausible patch is not enough. A retry can repeat an expired token, a conversion can erase leading zeros, and a pagination fix can silently lose records. API maintainers need a compact way to inspect what a proposed repair actually did.

What it does

API Repair Bench is an evidence-first review desk for synthetic API integration repairs. NVIDIA Nemotron proposes a narrow replacement function, Nebius Sandbox runs it remotely, and the review desk presents the generated source, contract checks and unresolved outcomes together.

The public demo is an interactive replay of recorded provider results, with downloadable evidence records. Three separately labelled reference demonstrations cover pagination, data conversion and token refresh. Browsing the demo never triggers a paid model call or executes generated code. The open-source CLI can reproduce fresh inference and Sandbox testing with the reviewer's own Nebius credentials and an explicit paid-action opt-in.

How we built it

Python supplies the challenge contracts, proposal validation, durable request reservations, remote runner and external comparison. HTML, CSS and JavaScript provide a responsive review interface. Optional WebMCP lets an agent select the same recorded results a person can inspect.

We use NVIDIA nvidia/Nemotron-3_5-Lightning through Nebius Token Factory and the Nebius Sandbox SDK. Token Factory removed the need to provision a GPU. Sandbox kept generated code off the development Mac. Only the original synthetic candidate, runner and job input are sent remotely; expected outputs remain outside the candidate runtime.

Model calls have explicit response allowances and no automatic retries. Source and contract hashes bind each proposal to its evaluation. Runtime output is bounded and retained. Completed failures can supply actual exception diagnostics for one correction; incomplete and uncertain attempts remain visible.

Built with AI assistance under founder review. This project contains original hackathon code and fictional fixtures, not private DeltaX engine code or customer data.

Results and challenges

An authentication baseline passed four of six checks, and its first correction still passed four. Our contract had omitted that responses were dictionaries. A separate proposal after making that interface explicit passed all six checks. We do not call that self-correction: the contract changed.

The amount-unit decision controls distinguish missing information from sufficient information. The model asked whether an undocumented amount used major or minor currency units; with an explicit major-unit specification, it chose to proceed. These are decision checks, not runtime validation of that control's patch.

Other attempts exposed truncated responses, structural rejection and a data-conversion error involving a nonexistent Decimal method. We kept those results and improved the feedback path to report bounded actual runtime exceptions. With actual runtime exception feedback, a fresh data-conversion baseline improved from four of nine checks to nine of nine after one correction on the same contract, preserving every previously passing check. Earlier unsuccessful correction pairs remain in the record. The final frozen evidence includes all 16 attempted model calls: 14 returned responses and two uncertain timeouts. Authentication and the corrected conversion each have a successful synthetic result; pagination did not yield a valid patch in this evaluation. Both amount-unit decision controls matched their expected decisions.

An SDK exit event also returned an integral float where our adapter expected an integer. We fixed the adapter and performed an explicitly recorded fresh verification, preserving the original inconclusive execution. That was an infrastructure bug, not a model failure.

What we learned

A useful repair workflow needs to distinguish incomplete specifications, broken patches and broken measurement. Test failures must tell the model what actually happened, while protecting private expected answers. The interface should make these distinctions clear to a reviewer without presenting a successful-looking response as proof.

Limits and next steps

This is a small synthetic evaluation, not a general reliability benchmark. Candidate code shares a process with the observation runner, so matching returned observations are not tamper-proof execution attestation. No generated changes are deployed automatically. Broader integration cases and stronger separation between execution and trusted measurement are next.

The judging demo is free to browse and requires no sign-in. The video is a narrated walkthrough of saved test results, clearly labelled as recorded.

Built With

Share this project:

Updates

Submission history