Inspiration
Every time I've given an agent write access to something that matters, I've had the same bad feeling. You get two options. Let it write, and hope. Or don't let it near anything, and do the work yourself.
Neither is good. The second one is why most "AI in the workflow" demos are really just chat with extra steps.
What I wanted was the thing code review already figured out years ago. You can open a pull request against a repo you can't push to. Nobody thinks that's a limitation. It's the whole reason the arrangement works: proposing and merging are different acts, and only one of them needs trust.
Floor plans turned out to be a good place to try it. They're shared, they're spatial, and a bad edit is obvious the moment you look at it. Nobody needs a diff explained to them when a wall has moved into a stairwell.
What it does
Onionskin is a multiplayer floor-plan editor built on WebMCP.
An agent connected to the page can do real work. Move walls, add doors, place desks, split rooms, check whether the plan still meets egress rules. What it can't do is change the plan. Every edit it makes lands on a forked layer. A human sees the diff, decides, and only then does anything reach the shared document.
The part I care about is that the guarantee is structural. Tools never receive the document. They get a narrow capability object that exposes the proposal layer and nothing else. A tool that tries to write to the committed plan doesn't get rejected at runtime, it fails to compile. There's a test that walks the whole source tree and fails the build if any file outside a short, explicit list so much as names the document type, so this stays true for tools nobody has written yet.
The tools also appear and disappear as you work. Select a wall and door tools show up. Open a proposal and begin_proposal withdraws, because you can't open two. There is deliberately no tool that commits. Not a tool that checks permissions and refuses. There just isn't one.
The plan syncs between browsers through a Cloudflare Durable Object, so two people, or a person and an agent, work on the same document at once.
How we built it
TypeScript throughout, strict, with no hand-written JavaScript anywhere in the app.
The document is a Loro CRDT, which is what makes forking a proposal cheap. A proposal is a fork; accepting is a merge; discarding is dropping the fork. Undo runs in separate lanes so a human's undo never silently reverses an agent's accepted work.
The WebMCP layer is document.modelContext.registerTool, built against the official webmcp-types package so the shape is checked by the compiler rather than by me squinting at a spec. WebMCP has no unregisterTool, so withdrawal happens by aborting the AbortSignal handed over at registration. Sync calls are queued, because selection can change faster than a registration round-trip resolves and two overlapping syncs could otherwise leave a tool registered that belongs to neither state.
341 tests. Every task went through an independent review before it counted as done, and that review found things I'd have shipped.
Challenges we ran into
The plan I wrote for myself was wrong more often than the code was. Five times an implementation stopped and reported that the reference code in my own plan had a defect, and it was right every time.
The one worth telling: a test in the suite was pinning a bug. It asserted that clicking "Accept and commit" with nothing accepted fires a commit with an empty set. That's exactly what shipped, and it meant a human clicking the obvious primary button silently destroyed the agent's staged work. No warning, no undo. I only found it by driving the deployed app in a browser. The tests were green the whole time, because one of them described the defect as the contract.
Then there's the pattern that showed up four separate times: correct mechanism, confident wrong justification. A comment arguing for why deduplication had to include values, twelve lines above code that discarded them. A test named as a backstop for a file it structurally excluded. A privacy rationale scoped to the wrong threat model. In every case the code survived scrutiny and the story about it didn't. I changed how I ran reviews after the third one, and the change was small: ask whether the stated reason is true, separately from whether the code works. That caught the last two.
The deploy that succeeded and served nothing was a good day too. Worker uploaded fine, 1.35 KiB, Durable Object live, and every path except the WebSocket returned 404 because I'd never configured static assets. The test asserting a non-room path "doesn't touch the namespace" passed happily. It just never checked the request went anywhere.
And the honest one: no stable browser ships WebMCP yet. Chrome needs 149 or later behind a flag. ChatGPT's in-app browser reads page text rather than calling tools, so it opened the site and asked me to upload a screenshot of the floor plan instead.
Accomplishments that we're proud of
The compile-time boundary is real, and I checked rather than assumed. Adding a write method to the capability type produces an actual compiler error, and a reviewer reproduced that independently before I believed it.
The agent console. Since there's no browser to watch driving the page, the app shows its own WebMCP surface: which tools are in scope right now, what you sent, what came back. Building it immediately exposed a bug in my own registry, which reported what the browser had accepted rather than what was in scope, so it showed an empty list in exactly the browsers it existed for.
And the eval suite reports skips as skips. Four of seven cases run without a live model; three don't. A green run that quietly skipped half its cases would be the worst thing to hand a judge.
What we learned
Tests can't see the browser. Both of the worst bugs, the deploy that served 404 and the button that destroyed work, sat behind a fully green suite. What they had in common is that each was guarded by a test asserting what the code didn't do rather than what it did.
Confident writing in a comment stops people checking it. That sounds obvious written down. It didn't feel obvious while I was reading past the fourth one.
The thing I didn't expect was how often "who is this actually for" broke a deadlock. What an agent should be told and what a human should be shown kept turning out to be two different answers, and every time I stopped trying to make one thing serve both, the design got simpler.
What's next for Onionskin
Verifying against a real WebMCP client the moment one is reachable. The page side is done and tested; the browser side is the last unverified link, and I'd rather say that plainly than claim a client I never tested.
After that, conflict resolution when two agents propose against the same objects. An agent-authored rationale attached to each proposal, so the human sees why and not just what. And replay that scrubs back through every proposal a document ever received, including the rejected ones, which is the part I actually want to look at.
Built With
- mcp
- typescript

Log in or sign up for Devpost to join the conversation.