An agent can fill the form. You hold the pen.

Hold the Pen is a benefits-style claim form that an AI agent can read, explain, and fill through WebMCP tools, while every value the agent writes stays marked, reviewable, and undoable, and no tool can submit the claim. Submission is deliberately absent from the WebMCP tool contract; review and declaration happen in the page, through controls no tool can reach.

Live: https://holdthepen.vercel.app · Repo: https://github.com/nischal94/holdthepen (MIT)

Why WebMCP is the right fit

Consequential forms (benefits, insurance, immigration) are where people most want help and least want to lose control. Hold the Pen is a benefits-style claim whose question wording models a real UK Universal Credit-style claim (income received vs earned, household definitions, carer thresholds); the programme, rules, and submission are fictional.

WebMCP lets the page hand an agent typed tools against the live form the person is looking at, instead of the agent guessing at the DOM. That is what makes a hard line possible: the page decides what an agent can do, and can make "submit" something no tool does. Seven tools are registered once on the document: get_claim_state, explain, review_agent_entries, fill_field, clear_field, navigate_to_section, prepare_submission_review. There is deliberately no commit tool.

How it makes the experience better

Three layers, three owners. Understand (agent, read-only): explain any question, term, or the consequence of each truthful answer, without recommending one. Fill (agent): write a value that is attributed, revision-checked, and undoable. Decide (person): review every agent entry, tick the declaration, submit. Those controls live in the page; no tool can reach them.

Every agent write shows a text badge, "Filled by the agent, not yet reviewed", lands in a persistent review queue with Accept / Correct / Clear, and is announced to screen-reader users by field name, never by value. The submit button stays disabled, and says why, until every agent entry is reviewed and the declaration is ticked. Any edit after the agent stages a review invalidates it. The page works with no agent at all, and a recorded demonstration shows the flow for browsers without WebMCP.

What a person and an agent can now do together

Before: an agent helping with a form either typed into fields with no trace of what it changed, or stopped at "you should fill this in yourself". Now: the person asks "explain the income question before I answer it" and gets the meaning; says "fill in the household section" and watches three fields fill with a visible mark and a queue count of 3; asks "get it ready for me to check" and the agent names the four missing required fields in form order and cannot go further. The page protects the person in two ways: while they are editing a field, an agent write is refused with CONFLICT_FOCUSED; once they have answered it, the agent cannot replace it and gets CONFLICT_HUMAN_VALUE. The recorded demonstration shows the second case. The agent prepares; the person approves. Verified end to end in the ChatGPT desktop browser on the deployed site.

How WebMCP was implemented

  • Imperative API only (document.modelContext.registerTool), because the ChatGPT browser does not support the declarative form API or iframe tools. Registration happens exactly once per page through a manager that reports a rejected registration as a visible degraded state and never unregisters (re-registering a changed schema is a race in the draft spec).
  • Tools read a framework-independent store at call time via getSnapshot(); React subscribes through useSyncExternalStore. A callback closing over React state would read the state at mount forever; a regression test writes, dispatches N times, then reads through the tool.
  • Every failure returns the same envelope (code, problem, cause, fix, retryable) so the agent self-corrects. Strict validation in code, loose in schema. Outputs capped at 1.5K characters. Read-only tools carry readOnlyHint; tools that echo user text carry untrustedContentHint. AbortSignal is honoured, and both toolcancel and toolcanceled spellings are handled.
  • 83 automated tests, including a typed fake of document.modelContext that enforces the browser contracts (duplicate names rejected, JSON-string arguments, cancellation events), axe in component tests, and eval fixtures in Chrome's expectedCall format for direct, ambiguous, and mid-chain prompts.
  • Honest limit, stated in the README: the page cannot tell a human click from an automated one. A trustworthy human-only approval boundary is a platform primitive WebMCP does not yet have (spec issues #165 and #277); prepare_submission_review is a stand-in for elicitation, not elicitation.

Built With

Share this project:

Updates

Submission history