-
-
Publish was approved and SPENT. The agent then tried to delete — refused, different fingerprint. Consent does not transfer.
-
The agent asked to publish. Nothing ran. The dialog names the exact action, its fingerprint, and the terms: this action, once, 120s.
-
The studio with four tools registered. Publish, share and delete do not exist yet — there is nothing to act on.
-
A guided flow: teach it, try it, put it live. The reply is composed from the knowledge you gave it, in the tone you chose.
Inspiration
We published a paper this year about trust in the Model Context Protocol — Whose Stranger Is It? (DOI 10.5281/zenodo.20847425). It argues for a deterministic guard: the agent proposes, the policy decides, and the execution happens outside the model.
WebMCP sharpened that argument in a way we did not expect. The standard carries
two annotations that look like safety metadata — readOnlyHint and
untrustedContentHint — and both are declared by the page itself, the very
party that may be the adversary. There is no signing, no revocation, no audit
levels. That is exactly where the paper's assumptions fail in a browser tab.
So we stopped writing about the guard and built one.
What it does
Assistant Studio is where you build the chatbot that answers your customers. You give it a purpose, teach it what it should know, try it the way a customer would, then put it live on a channel.
The whole studio works for a person alone, with no agent anywhere. WebMCP is an enhancement on top of a working product, never a precondition for it.
When an agent is present it drives the same studio through seven tools. Four run freely. Three — publish, share, delete — are seen by other people or cannot be undone, and those three do not execute on the agent's say-so.
Why this use case fits WebMCP. Building an assistant is long, iterative, unglamorous work: write a purpose, paste knowledge, test, notice the tone is wrong, fix, test again. An agent is good at that loop; a person has the taste to judge the result. What makes it the right demo is what happens at the end of the loop: the output gets published to customers, shared with colleagues, or deleted. The collaboration is genuinely useful and the boundary genuinely matters. Most agent demos have one or the other.
How it improves the experience. Before: four screens and about ten clicks per assistant. After: one sentence to the agent, and the interface moves in front of you — an assistant appears in the list, knowledge lands in its panel, a test reply comes back. You are not reading a transcript of what an agent claims it did; you are watching your own application change. The agent does not replace the interface, it shares it.
At the boundary the experience inverts on purpose. When the agent asks to
publish, the studio stops and shows the exact action in plain words — Publish
"Cafe Bot" to the whatsapp channel — composed by the application from the
arguments, never phrased by the model. The friction is the feature, and it
appears at precisely three places out of seven.
What is now possible that was not. An agent that builds another agent, with a person watching it happen live rather than reviewing it afterwards, and with a limit that holds regardless of what the model decides. Publish, share and delete cannot occur without a click made at that moment, on that action, with those arguments. One approval never becomes two. An approval for one assistant never moves to another. An approval to publish never becomes a deletion.
How we built it
Next.js App Router, TypeScript, Tailwind v4. State lives in memory: no database, no auth, no API keys in the repository. Test replies are composed deterministically from each assistant's own knowledge and tone, because the subject here is the consent gate, not the wording.
The WebMCP implementation. Seven tools registered through
document.modelContext.registerTool, each with a precise JSON Schema and enum
constraints wherever the value space is closed, so the model guesses less.
Lifecycle. Registration is bound to component lifetime with AbortSignal.
State-dependent tools. The three sensitive tools are registered only while at
least one assistant exists. They appear on the first create and unregister when
the last is deleted, emitting a real toolchange event. A tool with no valid
target does not exist.
The guard. Every execute is wrapped by guarded(). It reads sensitivity from
a build-time manifest — never from the arguments, never from the model — and for
sensitive tools it requires a consent bound to the tool name and to
SHA-256(tool ‖ NUL ‖ canonical(args)). Missing consent means the function is
never called, and the agent receives a refusal explaining that a dialog is now
open and that it cannot grant this itself.
The ledger. In memory, single-use, 120-second TTL, no localStorage: yesterday's
approval must not be usable today. grant has exactly one call site in the whole
repository — the onClick handler in the confirmation dialog. You can verify that
with a single grep, and the README tells you how.
Challenges we ran into
The WebMCP registry is a single global keyed by tool name, which means two concurrent registrations do not compose: the second overwrites the first, and then the first one's teardown unregisters the survivor. React StrictMode's double mount is exactly that race. It presented as tools that registered and then silently vanished. Registration is refcounted per document now.
A subtler one: useSyncExternalStore requires a referentially stable server
snapshot. An inline () => [] allocates a fresh array every render, and React
threw "Maximum update depth exceeded" on first paint. Frozen module-level
constants fixed it — and the same hazard was hiding in the consent ledger's
entries(), which rebuilt and re-sorted on every read.
Accomplishments that we're proud of
Eighteen tests, run with pnpm test. Four are the invariants from §7.4 of the
paper: a sensitive action without consent is dropped; the same action with
matching consent executes; consent for one action does not pass another; a
consumed consent does not replay.
Ten more assume a hostile model rather than a careless one: twenty hammered calls yield one dialog rather than twenty (flooding a person with modals until one is clicked is itself an attack), reordered argument keys cannot buy a second execution, expiry is checked at the millisecond on both sides, a guessed hash fails without burning a legitimate consent, and sensitivity passed as an argument changes nothing.
All four invariants were also verified end-to-end in a real browser against the live domain, not only at the unit level.
What we learned
Honestly, and this matters more than anything else in this submission: it does not solve prompt injection, and the paper it comes from says so explicitly. A compromised model still chooses which action to propose and how persuasively to propose it.
The guarantee is deliberately narrower and worth stating exactly: a sensitive action never happens silently, and consent is contemporaneous, single-use, and non-transferable. The guarantee holds outside the model, which is the only place it can hold.
We also learned that the interesting failures are not in the cryptography. They are in the plumbing — a global registry that silently drops registrations, a render loop that eats the answer you just asked for. The gate was the easy part.
What's next for Assistant Studio
Where the standard's page-declared annotations fail, and how an agent behaves when it meets a refusal it cannot talk its way past, becomes material for the next version of the paper. The repository stays public and MIT-licensed as a citable reference implementation.
Built With
- next.js
- react
- tailwindcss
- typescript
- vercel
- webmcp


Log in or sign up for Devpost to join the conversation.