-
-
The claim as it loads. Seven WebMCP tools are registered, and the header states who does what.
-
Your own entry stays yours, in green. The agent's entries are marked in amber and wait in the review list.
-
The recorded demonstration runs the same tools. The agent's write is refused because the person answered first.
-
Submit is disabled and says why: six agent entries are still unreviewed.
-
Submitted by the person, with a reference. No tool could have done this.
An agent can fill the form. You hold the pen.
Hold the Pen is a benefits-style claim form that an AI agent can read, explain, and fill through WebMCP tools, while every value the agent writes stays marked, reviewable, and undoable, and no tool can submit the claim. Submission is deliberately absent from the WebMCP tool contract; review and declaration happen in the page, through controls no tool can reach.
Live: https://holdthepen.vercel.app · Repo: https://github.com/nischal94/holdthepen (MIT)
Why WebMCP is the right fit
Consequential forms (benefits, insurance, immigration) are where people most want help and least want to lose control. Hold the Pen is a benefits-style claim whose question wording models a real UK Universal Credit-style claim (income received vs earned, household definitions, carer thresholds); the programme, rules, and submission are fictional.
WebMCP lets the page hand an agent typed tools against the live form the person is looking at, instead of the agent guessing at the DOM. That is what makes a hard line possible: the page decides what an agent can do, and can make "submit" something no tool does. Seven tools are registered once on the document: get_claim_state, explain, review_agent_entries, fill_field, clear_field, navigate_to_section, prepare_submission_review. There is deliberately no commit tool.
How it makes the experience better
Three layers, three owners. Understand (agent, read-only): explain any question, term, or the consequence of each truthful answer, without recommending one. Fill (agent): write a value that is attributed, revision-checked, and undoable. Decide (person): review every agent entry, tick the declaration, submit. Those controls live in the page; no tool can reach them.
Every agent write shows a text badge, "Filled by the agent, not yet reviewed", lands in a persistent review queue with Accept / Correct / Clear, and is announced to screen-reader users by field name, never by value. The submit button stays disabled, and says why, until every agent entry is reviewed and the declaration is ticked. Any edit after the agent stages a review invalidates it. The page works with no agent at all, and a recorded demonstration shows the flow for browsers without WebMCP.
What a person and an agent can now do together
Before: an agent helping with a form either typed into fields with no trace of what it changed, or stopped at "you should fill this in yourself". Now: the person asks "explain the income question before I answer it" and gets the meaning; says "fill in the household section" and watches three fields fill with a visible mark and a queue count of 3; asks "get it ready for me to check" and the agent names the four missing required fields in form order and cannot go further. The page protects the person in two ways: while they are editing a field, an agent write is refused with CONFLICT_FOCUSED; once they have answered it, the agent cannot replace it and gets CONFLICT_HUMAN_VALUE. The recorded demonstration shows the second case. The agent prepares; the person approves. Verified end to end in the ChatGPT desktop browser on the deployed site.
How WebMCP was implemented
- Imperative API only (
document.modelContext.registerTool), because the ChatGPT browser does not support the declarative form API or iframe tools. Registration happens exactly once per page through a manager that reports a rejected registration as a visible degraded state and never unregisters (re-registering a changed schema is a race in the draft spec). - Tools read a framework-independent store at call time via
getSnapshot(); React subscribes throughuseSyncExternalStore. A callback closing over React state would read the state at mount forever; a regression test writes, dispatches N times, then reads through the tool. - Every failure returns the same envelope (
code,problem,cause,fix,retryable) so the agent self-corrects. Strict validation in code, loose in schema. Outputs capped at 1.5K characters. Read-only tools carryreadOnlyHint; tools that echo user text carryuntrustedContentHint.AbortSignalis honoured, and bothtoolcancelandtoolcanceledspellings are handled. - 83 automated tests, including a typed fake of
document.modelContextthat enforces the browser contracts (duplicate names rejected, JSON-string arguments, cancellation events), axe in component tests, and eval fixtures in Chrome'sexpectedCallformat for direct, ambiguous, and mid-chain prompts. - Honest limit, stated in the README: the page cannot tell a human click from an automated one. A trustworthy human-only approval boundary is a platform primitive WebMCP does not yet have (spec issues #165 and #277);
prepare_submission_reviewis a stand-in for elicitation, not elicitation.
Built With
- next.js
- react
- tailwind-css
- typescript
- vercel
- vitest
- webmcp
Log in or sign up for Devpost to join the conversation.