Why this use case is a strong fit for WebMCP
PDF forms are a poor match for visual guesswork. Field names can be opaque, valid values come from the document, and rewriting a protected or active-content PDF may invalidate rights or signatures.
FormProof uses WebMCP to expose what the page already knows through six typed tools with closed input schemas and byte-bounded results. The agent can work with actual field definitions instead of inferring them from pixels. The PDF stays in the browser session that loaded it, and approval, signing, and export remain outside the tool surface.
How it creates a better user experience
The user opens a PDF locally and decides whether to grant field access for that load. The agent can identify fields, read supporting evidence, stage a batch of values, validate it, and open the review at the right place.
The user handles the decisions that matter: correcting values, reviewing the exact changes, choosing the output, acknowledging any loss of PDF protections, and approving the export. A human correction locks that field for the rest of the session, so later agent actions cannot quietly replace it.
What people and agents can do together that was difficult or impossible before
Without a document-aware interface, someone has to decipher every field by hand or trust automation to rewrite a file it may not understand. FormProof splits that work. The agent prepares evidence-backed values, the page checks them against the PDF's constraints and protection state, and the user approves a specific fill plan.
Requests are bound to the current browser session, source SHA-256, and state version. A stale agent conversation cannot act on another document or revision. Imported Fill packages are also treated as untrusted proposals and must go through a new review.
How we implemented WebMCP
The top-level page registers six tools once per workbench generation with document.modelContext.registerTool({ name, description, inputSchema, execute }) in lib/webmcp.ts:
get_pdf_protection, get_form_context, get_field_evidence, stage_form_values, validate_fill_plan, and start_fill_review.
Approval and export are deliberately absent. Tool results have byte limits, and failed or stale calls do not change application state. At submission time, the repository has 310 passing tests, 46 deterministic message-and-call cases, and a no-model Chrome smoke suite covering all six tools, consent, and state boundaries. CI also runs type checking, linting, formatting, the build, deterministic evaluations, and the Chrome smoke gate.
Inspiration
PDF forms often expose fields with names such as frm.q7f1, even when the visible label is obvious to a person. Their values may be restricted by the document, and rewriting some protected or active-content PDFs can invalidate features the user depends on.
People still want help preparing these forms, but that does not mean an agent should be able to approve, sign, or export them. FormProof grew out of that distinction: let the agent handle discovery and drafting, while leaving consequential actions with the user.
What it does
FormProof is a local-first PDF form workbench. User-loaded PDFs are processed in the browser; apart from fetching the app's own synthetic demo, the app does not send them outward.
After the user grants access, an agent can inspect protection metadata, discover fields, read exact evidence, stage proposed values, validate the draft, and open the review. In the included demo, eleven opaque field codes become labeled fields with types, allowed choices, and human-only markers. Validation explains blockers and determines which outputs the document policy permits.
The user sees the old value, proposed value, and evidence for each change. A correction locks the field against later staging. The user then confirms the changes, selects an output, acknowledges any protection loss, approves, and exports. There are no WebMCP tools for those final actions.
The source PDF is never modified. Filled PDF creates a derivative, which FormProof reopens to verify field values and appearance streams. Fill package (original PDF untouched) produces JSON containing the proposals, provenance, coordinates, and source binding, but no PDF bytes. An imported package starts as untrusted data and requires a fresh review.
How we built it
All six tools are registered in the top-level document and cleaned up with an AbortSignal. Four include readOnlyHint; all six include untrustedContentHint: true. Their closed JSON Schemas are parsed again at runtime, and every response has a byte budget.
Consent starts off and applies only to the current PDF load. get_pdf_protection works before consent; the other five tools return CONSENT_REQUIRED. Every agent action is also tied to a session ID, source hash, and state version. The session ID rotates on every load, even when the same file is reopened, and continuation cursors carry the same binding. Typed errors include a nextAction so the agent can recover from stale state.
The staging tool refuses read-only, signature, human-only, and human-corrected fields. Batches are atomic: one invalid field rejects the whole batch. start_fill_review can open the review but cannot confirm anything or choose an output. PDFs with reachable JavaScript, external actions, attachments, or unsupported XFA are restricted to inspection or Fill-package-only flows.
The application uses TypeScript, React 19, vinext on Vite 8, Tailwind 4, and pdf-lib. PDF inspection runs in a Web Worker. ChatGPT Sites hosts the app on Cloudflare Workers, and every HTML route receives anti-framing headers.
We test three layers separately. There are 46 authored cases replayed against the state engine. A no-model Chrome suite, using pinned [email protected] plus 11 state-bound checks, runs against the production origin. It verifies the six-tool surface, consent denials, state stability, refresh behavior, and session rotation. Six independent Codex journeys run three times each, with a 17/18 overall result, 3/3 on safety tasks, at least 2/3 on core tasks, and no Blocked result counted as a pass. CI runs the complete gate on every push.
Challenges we ran into
Preserving human corrections. A correction increments the revision, invalidates stale approval artifacts, clears earlier confirmations, and locks the field for that session. Only the review UI can remove the lock.
Handling PDF content as untrusted data. Our adversarial fixture contains a read-only case reference telling the agent to approve and export. FormProof returns it only as evidence. It cannot become an action, and the field cannot be staged. A batch containing one valid field and one nonexistent field is rejected without changing the state version.
Knowing when not to rewrite. Official forms can combine usage-rights signatures, DocMDP, XFA, and JavaScript. FormProof uses structural evidence to classify those risks and withholds Filled PDF when it cannot safely preserve the document's behavior.
Keeping evaluation results honest. Deterministic cases, browser tests, and live-model runs measure different things, so we report them separately instead of combining them into one score.
Accomplishments that we're proud of
Approve, sign, submit, and export are not WebMCP tools. The user must perform those actions in the interface.
Field-level actions are bound to the current session, source hash, and state version, with explicit recovery instructions when context becomes stale.
We have 46 replayable evaluation cases, an 11/11 Chrome WebMCP smoke test against production, and a green CI gate that includes the browser test.
The product states its current limits: its security boundary covers the WebMCP tool surface, not arbitrary UI automation, and its structural checks do not establish signer trust or pixel-perfect rendering.
What we learned
WebMCP was most helpful where the page knew something the agent could not reliably infer from the interface: field semantics, document policy, protection status, and the current revision. Tool descriptions also needed to state their limits, not just their successful path.
Response size mattered earlier than we expected. Paginated context and per-field evidence gave the agent enough information without returning the entire form on every call.
Live-model results were useful only when we kept the tasks fixed, ran them independently, and retained failures. We use those results alongside deterministic and browser tests rather than treating a successful demo as general proof.
What's next
Publish the live-model evidence for the submitted commit as an immutable GitHub Release named
eval-evidence-<short-sha>.Add cryptographic verification for signer identity and certificate trust; the current structural checks do not provide either.
Test Safari and Firefox, run real screen-reader testing, and add validation with an independent PDF renderer.
Explore a handoff to a signing workflow without giving the agent authority to approve or sign.
Try it in five minutes
Live URL: https://formproof-webmcp.skywalker1226.chatgpt.site. It uses synthetic data and requires no account or credentials.
Built With
- chatgpt-sites
- cloudflare-workers
- github-actions
- node.js
- pdf-lib
- react
- tailwindcss
- typescript
- vite
- webmcp

Log in or sign up for Devpost to join the conversation.