Inspiration

WebMCP lets an agent take useful actions inside a website. But the person does not stop using the app while the agent works.

That creates a quiet and dangerous race: an agent reads the current state and starts a slow task. The person makes a newer change. The delayed agent action then finishes with old information and overwrites the person's work. Every request can still report success, so normal logs may never explain what went wrong.

We built Interleave to catch that moment, explain it clearly, and turn it into a test that prevents the same bug from returning.

What it does

Interleave records one shared session across the agent, the person, and the application. It captures the pending tool call, meaningful human edits, state revisions, and the final write on one ordered timeline.

When a rule is broken, Interleave shows:

  • the exact product rule that failed;
  • what the final value should have been;
  • what value was actually written;
  • which state revision the agent started from;
  • and which newer revision existed when the delayed work finished.

The developer can then replay the same recording, compare the current behavior with a proposed fix, remove every unnecessary step, and export a deterministic regression test.

Our flagship demo uses public issue makeplane/plane#9674. Plane queues a metadata crawler. A person saves a newer title while that crawler is waiting. The delayed worker later writes an older title over the human edit. Interleave witnesses the overwrite, reduces it to three required actions, and replays the same recording against a revision guard that preserves the person's work.

The workbench also includes reservation and TodoMVC adapters to show that the recorder is reusable across different application models.

Why WebMCP is essential

Interleave is built around a real WebMCP interaction, not a chatbot wrapped around a dashboard.

The Plane demo exposes nine native tools while idle. The agent can read the incident, start the delayed update, replay a recording, compare implementations, reduce a failure, and export proof. While the crawler is pending, the tool surface changes to five controls that match the application's live state: read, save newer metadata, hold, complete, or cancel.

The slow tool returns a real pending promise. While that promise is unresolved, the person can keep using the visible application. That shared moment is the core experience: WebMCP gives the agent a clear semantic action, while Interleave observes how that action interleaves with human work.

We use document.modelContext.registerTool directly. There is no WebMCP polyfill and no reconstructed agent transcript. Tools have clear names, input descriptions, state-aware availability, AbortSignal cleanup, and compact evidence receipts. Large recordings, tests, and patches stay behind explicit downloads instead of filling the agent's context.

A better experience for people and agents

Without Interleave, a person can watch a correct edit disappear and have no explanation. With Interleave, the app can show exactly which delayed action caused the loss, what should have happened, and whether a guard fixes it.

This also changes what developers can do. A hard-to-reproduce session becomes a small, repeatable test. A person can continue working while an agent performs a slow task, and both can verify that the latest human intent survived.

How we built it

The application uses Next.js, React, TypeScript, Tailwind CSS, and native WebMCP. The reusable @interleave/recorder package has no runtime dependencies and no React dependency.

Each integration supplies four things: a focused state reader, meaningful human actions, a correctness rule, and replay behavior. The recorder handles calls, results, errors, cancellation, timestamps, before-and-after state, immutable snapshots, redaction hooks, interruption, and bounded history.

The Plane integration is pinned to public commit da1a7ab. We traced the issue from the API handler to the queued worker, built a deterministic local model, prepared a compare-and-set patch, and added focused worker and endpoint contract tests. The fixture never contacts a live Plane deployment or user data.

Challenges we ran into

Keeping the race real

An animated log would look convincing but prove very little. We had to keep one actual tool promise pending while a real human edit changed the same state, then record the delayed completion at the correct checkpoint.

Changing tools without cancelling the active call

The available WebMCP tools need to change while work is pending. Re-registering too early can clean up and abort the call being recorded. Interleave waits for the active call to settle before replacing the tool surface.

Recording meaning instead of noise

Raw DOM mutation streams are large and fragile. Interleave records semantic actions such as edit_metadata and add_todo, so the person's intent survives in the evidence even when the stale write erases it from the screen.

Reducing an asynchronous failure

Reduction must rerun each candidate against fresh state. Interleave removes one command at a time and keeps it removed only when the same rule still fails.

Accomplishments we are proud of

  • A complete native WebMCP acceptance run in the browser: dispatch, real human click, delayed overwrite, replay, comparison, reduction, and export.
  • A three-command minimal reproduction for a documented open-source bug.
  • A proposed Plane patch with all 37 tests passing in the two affected upstream modules, including 13 new cases.
  • A 50-test Interleave release gate covering the recorder, all three adapters, replay, reduction, exports, and the state-aware WebMCP surface.
  • A dependency-free recorder package that another WebMCP application can install and adapt.

What we learned

WebMCP can be more than a way for agents to click fewer buttons. It gives web applications a structured boundary where agent intent, human intent, and application state can be observed together.

We also learned that correctness rules must stay close to the product. Interleave provides the recording and proof system, while each application defines what must be preserved.

What is next for Interleave

Next we want to publish the recorder to npm, add adapters for editors and collaborative work tools, connect recordings to CI, and make browser-level regressions exportable to more test runners. We also want to take the Plane patch through upstream review and use that process to improve Interleave's contribution workflow.

Built With

Share this project:

Updates

Submission history