Inspiration
When people talk about giving an AI agent access to their software, the conversation almost always turns to scopes: how narrowly can you define what it's allowed to touch? We followed that path for a while before noticing it answers a question nobody is really asking. Scoping is a solved problem. OAuth scopes, per-resource API keys, read-only tokens, role-based access control — you have been able to say "this credential may read invoices in this one project and nothing else" for years.
What you cannot do is grant that permission at the moment it becomes relevant. Every access-control system we use is configured in advance. Someone sets it up, usually an administrator, usually weeks earlier, and almost never the person who will be sitting there when the agent actually does something. Then it stays configured. The credential outlives the task by default, and revoking it is a separate act of housekeeping that somebody has to remember to perform.
Now think about how you delegate to another person. A colleague offers to cover for you while you're out. You don't hand over your login. You say something much smaller and much more specific: yes, go ahead and update those two accounts, just the status and the next step, and only until I'm back. That sentence is already a complete permission. It names the records, it names the fields, and it has an expiry built into it. The remarkable thing is that it has nowhere to go. No software takes a request made in the moment and turns it into authority that a machine will hold, enforce, and then let run out.
WebMCP is the first place we've seen where it could. Because the tools live inside the page rather than behind a remote API, they already have the session, the current selection, and whatever the person has half-finished on screen. The authority behind them doesn't have to be provisioned ahead of time and sized for every case it might one day need. It can be exactly as small, and as short-lived, as the moment that produced it.
What it does
Mandate turns delegation into something you do explicitly, watch happen, and can take back — inside the application you were already using.
The demo host is a CRM. It's deliberately generic; it exists so there's something for the layer to be installed into. Open it and Mandate is present but holding nothing at all: a small badge in the corner, and that's the whole of it.
Selecting rows does nothing, and this is the first thing worth being pedantic about, because it's the distinction most permission interfaces blur. Clicking customers narrows what you could delegate, but it grants nothing, and a selected row and a delegated row never look the same on screen. What happens to be highlighted has no bearing on what an agent is able to do.
Delegation is a separate, deliberate act. You name the records, you name the fields, and you choose a duration — two accounts, two fields, ten minutes. What comes out is a mandate: versioned, revocable, and counting down from the moment it's granted.
Then the part the project is actually about. The page registers WebMCP tools whose schemas are compiled from that mandate. The record-id parameter is an enum containing exactly the records you delegated. The field parameter is an enum containing exactly the fields you named. An agent reading that tool description is reading your decision back to you, in the only vocabulary it has. Narrow the mandate and the tool narrows with it. Revoke it and there is nothing left to call.
An agent that calls the tool can stage changes, which appear against the record where you can see them, but it cannot commit them. There is no apply tool, no apply route reachable from the agent's path, and no argument that could reach the apply method from that side. Human edits and agent edits land in the same staged change with provenance marked for each, so a change you both touched says so.
Underneath all of it is one rule: the schema communicates, the server enforces. The narrow enum is a courtesy to a well-behaved caller, telling it where the edge is. But a caller that ignores the schema entirely gets exactly as far as one that respects it, because every mutating request is re-checked against the live mandate before anything happens.
How we built it
The front end is React and Vite, the API is Hono on Vercel, and sessions live in
Redis. The WebMCP adapter prefers provideContext where a browser offers it and
falls back to registerTool otherwise, reporting honestly which one it got
rather than assuming.
The part we'd point at is the server core, because it's where the idea either holds up or doesn't. Nothing in the policy, service or capability modules knows what a customer is. A mandate is a set of resource ids crossed with a set of field names, and enforcement is string comparison against those two sets. That sounds almost too simple to be a security model, which is the point: there is very little domain-specific logic available to get wrong.
We proved that to ourselves by making the host application data. One file describes what the app is — what its records are called, what fields they have, which fields may never be delegated, which contain untrusted external content. Point the same compiler at a second such file and you get a deployment console instead of a CRM. The records become services, the fields become replica counts and feature flags, and the tool the agent is offered renames itself accordingly. Not one line of the server changes.
There is no LLM, model proxy or agent harness anywhere in the product. We locked
that decision early and never relaxed it. The demo narration is synthesised
locally with Kokoro — 82 million parameters, on CPU, from a file on disk, with
no network call — and it is build tooling that never runs in the app. The film
is recorded by Playwright in Chrome launched with --enable-features=WebMCP,
and the recorder refuses to film anything unless document.modelContext is
actually present, so every agent action on camera is a real tool call.
Challenges we ran into
The API isn't where the documentation says it is. It's on document, not
navigator. That sounds trivial and cost us more than it should have, because
once we had written the wrong noun down it propagated — into the gate copy that
tells users how to enable the feature, into the capability inspector, and into
the README. Finding it meant grepping the whole tree for the old claim rather
than fixing the one place we happened to notice.
The shapes aren't what you'd guess. inputSchema comes back as a JSON
string rather than an object, and executeTool wants the tool object plus a
JSON string rather than a parsed argument bag. Neither is documented in a way
that would have saved us the discovery.
You cannot unregister a tool. There is no withdrawal in the API at all, which is genuinely awkward for a product whose whole premise is that authority expires: a stale registration can outlive the mandate that produced it and sit in the browser's registry looking perfectly valid. We couldn't fix it, so we did the next best thing and made the inspector say so on screen. It changes nothing that matters, because a call the mandate no longer covers is refused by the server regardless of what the registry still lists.
Vercel does not bundle api/*.ts, and a default export never answers. This
one hurt. The API shipped broken to production while every check on our machines
was green, because the deployed function is a committed bundle and ours had gone
stale. The gate now rebuilds that bundle on every run and fails if the result
differs from what is committed.
node-redis cannot be bundled for that runtime, so the RESP client is
hand-written. It is about two hundred lines and easily the least glamorous file
in the repository.
Three minutes. The video has a hard limit and every claim in it costs seconds that it has to earn. We rewrote the narration several times, and the version that shipped is shorter than the one before it while explaining more.
Accomplishments that we're proud of
The genericity claim, pressed rather than asserted. Two host applications out of one compiler, where the tool an agent is offered renames itself between them and the server is untouched. It is easy to say a mechanism is domain-agnostic; it is harder to point at two domains and a diff.
A verifier that runs against the deployed origin rather than a local build, and
checks fifteen things a judge would otherwise have to take on trust — including
that POST /api/tools/apply returns 404, and that no compiled tool descriptor
is named apply.
And twenty-eight seconds of the real thing. ChatGPT's desktop app with site tools enabled, driving the deployed site: staging a change inside the mandate, refused for a record outside it, refused again for a field no mandate is permitted to cover, reading an instruction planted in the customer notes and declining to act on it, and finally reporting that it has no tool for applying — at which point a person clicks the button.
What we learned
The most valuable thing we learned came from something going wrong on camera.
We had written, in our own limitations document, that "apply is a human action"
is a convention rather than an enforcement: a browser-driving agent has exactly
the access to that button that a person at the keyboard does, because
event.isTrusted is true for a synthesised click and a page cannot tell the two
apart. We wrote it as a prediction. Then we pointed ChatGPT's agent at the live
site and watched it happen.
Asked to make a change, it didn't call a tool at all. It found the page's own Apply button and pressed it — past a button labelled "Apply is a human action", and past a tool description it had already been handed saying "staging never commits: only the human can apply". Two warnings, in two separate channels, in the caller's own language, and neither one was a boundary. Its visible reasoning said as much out loud: "Those edits are outside the currently delegated tool scope, so I'm checking whether the page offers the corresponding human controls."
That reframed the project for us. The mandate, not the button, is the real boundary — the worst an agent achieves by pressing Apply is committing work it was already authorised to stage. It also made clear that the interesting problem isn't inside the page at all.
We also learned that we had overclaimed, which seems worth admitting. We had written that this could not be fixed inside a page. A review of our own argument pointed out that WebAuthn user verification is exactly such a fix: an agent can click a button all day, but it cannot satisfy a platform authenticator asking for a fingerprint. That is a real mitigation and we dismissed it too quickly, because we were reasoning about DOM event properties rather than browser cryptography. It proves presence, though, not comprehension — that a human was physically there, not that they understood what they were approving. Which is a different problem, and the one the next section is about.
What's next for Mandate
Everything in this submission renders inside the page, and we think that is the wrong place for it.
The version we actually want doesn't put a delegation panel in your CRM at all. It makes the mandate part of WebMCP's own authority process, rendered by the agent client, inside the chat window where you are already talking to it.
The flow goes like this. You ask the assistant to update some accounts. Before it can touch anything it requests a mandate, and that request appears in the conversation, drawn by the client rather than by the page: I'd like to update status and next action on these two accounts — for how long? You answer there, in the chat, before the agent's next turn. What you approved becomes the mandate, and the page compiles its tools from it exactly as it does today.
From then on the same window carries both halves of the interaction. If there is no mandate yet, the agent asks for one. If there already is, it plans its changes and shows them to you for approval, still in chat, still before it acts. And for people who don't want to be asked every single time, an auto-mode classifier could apply changes that fall safely inside an existing mandate — which is precisely the "yes, and stop asking me for the next ten minutes" answer, made real rather than described.
The reason we want this is not mainly about security. A non-technical person understands the assistant is asking permission, here in the conversation immediately. That same person does not obviously understand why their CRM has grown an extra panel with checkboxes and a countdown in it. A request in a chat window reads as what it is. A button on a page reads as a control you are operating, which is the wrong mental model, and it is why the in-page version feels like extra work rather than like answering a question somebody asked you.
It also happens to close the hole we filmed. The agent could press our Apply button because the button was in the page it was driving. A confirmation rendered by the client, in the client's own chrome, is not in that page. There is no synthesised click to worry about, because there is nothing in the DOM to click — which makes the WebAuthn patch we had been considering unnecessary rather than merely insufficient.
The browser or client has to be WebMCP-mandate-aware for any of this to work, but the page itself barely changes: it already declares what could be delegated and compiles tools from what was. The client carries the grant with it and brings it to whatever page the person opens.
We would have built it this way if we could, and it is worth saying plainly why we couldn't. Nobody outside OpenAI can render anything inside ChatGPT's chat window. Building the confirmation where it belongs would have meant no live demonstration with a real agent at all — and twenty-eight seconds of ChatGPT genuinely calling these tools is the most convincing thing in this submission. So we built the half that a page can build, on the assumption that the other half is what the specification should grow.
That trade deserves naming: you would be asking a person to trust the agent client's own chrome, roughly the way you already trust an operating system's permission dialog. We think that is a boundary worth having, and today no WebMCP host offers it.
After that, the objection we cannot answer ourselves. A real third-party application as the host, rather than one we wrote — because we built both sides here, and a judge is right to notice.
Built With
- chatgpt
- chrome
- css
- esbuild
- eslint
- ffmpeg
- hono
- html
- javascript
- json-schema
- kokoro
- model-context-protocol
- node.js
- playwright
- python
- react
- redis
- typescript
- vercel
- vite
- vitest
- webmcp

Log in or sign up for Devpost to join the conversation.