Inspiration

Software has dependency graphs, diffs, and CI. Research papers still copy results by hand.

A run changes. A configuration changes. A new figure replaces the old one. But the manuscript often stays frozen until someone notices the mismatch manually.

Potter's Wheel keeps the paper connected to the research that produced it.

What it does

Potter's Wheel is a provenance-aware research dependency system built around one continuous lifecycle:

Experiments → Evidence → Paper

Experiments are first-class objects with planned/running/completed/failed/superseded states, methods, parameters, scalar metrics, metric histories, datasets, artifacts, notes, source commits, timestamps, and run-to-run supersession.

Results can be published as typed evidence and linked directly to the manuscript objects that depend on them: claims, methods, tables, and figures. The trace is bidirectional: run → evidence → paper, and paper → evidence → run.

When newer evidence replaces older evidence, Potter's Wheel can mark downstream manuscript objects stale and generate a reviewable Research Diff instead of silently rewriting the paper.

Three signature experiences make that visible:

  • Research X-Ray turns the manuscript into an integrity map, highlighting verified, stale, contradicted, unlinked, and review-required objects directly in the editor.
  • Verify This deterministically checks a selected quantitative claim against its linked evidence, including percentages, decimals, grouped numbers, and scientific notation, while surfacing ambiguous or conflicting evidence for review instead of guessing.
  • Research Diff shows Git-like drift between the current research and the manuscript across results, methods/configs, tables, figures, and artifact identity. Corrections are staged visibly and require researcher acceptance.

The demo deliberately contains multiple kinds of drift: a paper that says 18.2% while the current experiment says 16.8%, a learning-rate mismatch, a stale table, and a stale figure.

Potter's Wheel also includes an agent-side experiment sync workflow: when a research agent starts, evaluates, or finishes a run, it can synchronize the visible Experiments and Evidence workspaces through WebMCP—recording parameters, metric history, artifacts, run status, evidence, and manuscript links without scraping or clicking the UI.

Why WebMCP

Potter's Wheel is fully usable without AI. A researcher can record and compare runs, publish evidence, write the paper, inspect provenance, verify claims, review diffs, and accept or reject changes manually.

WebMCP adds a second native control surface for the user's existing research agent.

The application currently registers 37 WebMCP tools spanning live manuscript context, selections, claims, figures, tables, provenance, integrity state, experiment reads/comparison, run creation and updates, metric logging, parameter logging, artifacts, run status and supersession, evidence publishing, evidence linking, verification, navigation, comments, and reviewable corrections.

That matters because the user's agent may already know the repository, experiment history, code paths, and research decisions, while Potter's Wheel knows the exact manuscript, selection, evidence graph, and integrity state currently open in the browser. WebMCP lets those two contexts meet in the same live workspace.

The agent can maintain runs, publish and compare evidence, verify claims, navigate stale dependencies, add comments, and stage corrections. It cannot accept its own manuscript rewrite. Final authority stays with the researcher.

How we built it

Potter's Wheel builds on Antmicro's Apache-2.0 licensed MyST Editor for the scientific editing foundation.

For this challenge, we added the experiment/evidence lifecycle workspace, provenance graph, typed evidence model, Research Integrity panel, Research X-Ray, deterministic quantitative verification, evidence-driven Research Diff, supersession logic, a reproducible research fixture, agent experiment-sync workflow, and the 37-tool WebMCP surface.

WebMCP tools operate on the same state visible to the human, use explicit input schemas, distinguish read-only from state-changing actions, mark researcher-authored content as untrusted input, return structured failures for stale or invalid context, and unregister with the editor lifecycle.

The browser suite currently passes 112/112 Playwright tests, covering the research workspace, provenance/integrity flows, WebMCP registration and actions, experiment lifecycle, evidence publishing, manuscript linking, stale-context guards, and review-only writes.

No account, model API key, or backend is required to run the deterministic judge demo.

Challenges

The hardest problem was not making the agent powerful; it was making it powerful without making it authoritative.

Research facts need stronger guarantees than generated prose. We therefore keep experiments, evidence, verification state, and manuscript dependencies structured and inspectable. Agent conclusions do not directly rewrite the paper. Changes are staged into the same review surface the researcher uses.

A second challenge was treating WebMCP as a real application interface rather than a demo wrapper. The tool surface needed stale-state protection, evidence gating, structured failures, safe navigation, lifecycle cleanup, and shared UI/tool state while the WebMCP API itself is still evolving.

What we learned

The most interesting part of the agent-native web is not adding chat to an application.

It is letting a user's existing agent enter a specialized tool with a first-class semantic control surface, bringing the context it already has while operating the exact same state as the human.

For research, that creates a new possibility: the agent that helped run the work can help keep the paper honest about the work.

What's next

The next step is turning Potter's Wheel from a self-contained research workspace into a professional research-integrity layer across the existing stack:

  • adapters for MLflow, Weights & Biases, DVC, and other experiment trackers
  • persistent multi-user provenance and collaboration
  • typed live evidence bindings inside manuscripts
  • a full provenance/dependency graph explorer
  • Research CI checks for stale or unsupported manuscript claims
  • immutable submission snapshots that freeze the exact runs, evidence, commits, datasets, figures, and manuscript version behind a paper
  • deeper Git/Overleaf-style interoperability

The long-term goal is simple: research artifacts should evolve without the paper silently drifting away from them.

Built With

Share this project:

Updates

Submission history