Inspiration

Browser agents can inspect and act quickly, but important product decisions should not collapse into “the model suggested it, so the model changed it.” Accessibility and design-system teams need a shared surface where an agent can use the live page and prior decisions while an exact visible UI review remains the only source of application authority.

Focus Contract Studio explores that future of the open web through one deliberately narrow, testable decision: which control should receive initial focus in a destructive Delete Account dialog.

What it does

Focus Contract Studio gives a reviewer and ChatGPT one shared, page-bound change-control loop. R13 makes that contract legible at first glance: a forensic editorial workspace shows the concrete Delete-versus-Cancel mismatch, then maps the Browser → Agent → Reviewer → Browser protocol before the first action.

The browser renders implemented revision 1, where the dialog initially focuses Delete. The page retrieves synthetic, applicable precedent saying this exact destructive-dialog context should initially focus Cancel, and reports a bounded DECISION MISMATCH. ChatGPT reads the live review through WebMCP and stages an immutable Delete → Cancel proposal. The renderer stays on revision 1 and the proposal stays visibly NOT APPLIED. The reviewer inspects the exact diff, digest, and base revision. There is no approval tool: only the visible UI-mediated decision can authorize application. ChatGPT applies that exact approved payload. Guarded writes create revision 2 and a durable receipt. The person completes a keyboard rehearsal in the real dialog. ChatGPT reads the page again, receives the exact current committed browser rehearsal, and verifies six behaviors: initial focus, focus order, forward wrap, backward wrap, Escape, and returned focus.

History, stale-state rejection, idempotent retries, reload recovery, revisioned undo, privacy limits, and a complete keyboard-accessible fallback remain visible on the same page.

Why this is a strong fit for WebMCP

Every operation depends on live page context: the rendered variant, implemented revision, browser observation, eligible precedent, immutable proposal, current review state, and committed rehearsal. A detached chatbot or generic API would lose the precise state boundary that makes each action useful and safe.

The page registers exactly four typed top-level tools with document.modelContext.registerTool(...):

read_active_focus_review create_focus_contract_proposal apply_approved_focus_contract verify_focus_contract

The WebMCP adapters and visible UI call the same protected application services. There is deliberately no approval, review, rehearsal-capture, reset, undo, URL, selector, or workspace-selection tool. Page text, retrieved precedent, tool output, and model language are untrusted evidence—not authority.

How it creates a better experience

The reviewer and agent work from the same live revision instead of exchanging screenshots and stale summaries. The agent can create a durable, inspectable proposal without silently changing the product. The exact visible UI review remains the only application authority. Application fails closed on stale revisions, changed digests, revoked or rejected decisions, foreign workspaces, and conflicting retries. Verification comes from fresh browser events, not from regenerating the expected answer from the configuration under test. The full visible workflow still works when WebMCP is unavailable.

Before this, teams could combine chat, tickets, component previews, precedent documents, and browser testing, but the state and authority boundaries lived in separate tools. Focus Contract Studio turns them into one legible human-agent protocol on the open web.

What people and agents can do together now

ChatGPT can interpret bounded precedent, create a field-supported proposal, apply only the exact payload already authorized in the page, and return a durable verification receipt. The reviewer can inspect every transition, supply the only approval authority through the visible UI, replay receipts, recover from uncertain responses, and undo a committed revision without trusting hidden agent state.

The current release retains the final page-bound handoff: after the browser rehearsal, read_active_focus_review returns a bounded verificationTarget for only the active workspace, variant, and revision. Foreign, stale, test-only, recording, uncommitted, and expired rehearsals are excluded. Judges no longer need a database console, copied cookie, workspace ID, or manually supplied rehearsal ID.

How we built it

Focus Contract Studio is a full-stack ChatGPT Site written in strict TypeScript with React and Next.js-compatible Vinext output. Sites-managed D1 stores isolated anonymous workspaces, focus revisions, synthetic precedent, immutable proposals, visible review decisions, allowlisted raw rehearsal events, and receipts.

Zod defines strict route and WebMCP schemas. One indexed D1 query filters precedent by workspace, product, component, behavior, status, and valid time before deterministic TypeScript BM25, structured applicability, relationship ranking, and Reciprocal Rank Fusion combine the eligible lists.

Proposal payloads use canonical SHA-256 digests. Apply uses expected revisions, scoped idempotency, repeated conditional guards, database constraints, and explicit zero-row handling. Verification consumes immutable finalized browser events rather than generating expected events from the configuration under test. The Site calls no model API: ChatGPT supplies reasoning through WebMCP, while deterministic code owns validation, persistence, authorization checks, application, and verification.

Challenges we solved

Preserving useful precedent without turning retrieved text into permission. Making proposed, reviewed, applied, and verified visibly and mechanically different states. Keeping retries, stale revisions, and concurrent applies from creating partial or duplicate mutation. Verifying the rendered dialog from raw events without recording optional reason-field text. Supporting an anonymous public judge flow with isolated, expiring, abuse-bounded workspaces. Making the final verifier target page-bound without introducing an unsafe implicit “latest rehearsal” mutation.

Accomplishments verified for R13

The deployed R13 source, release/webmcp-challenge-2026-r13 branch, and annotated webmcp-challenge-2026-r13 tag resolve to commit 3a37d92cb22d39602acbb3bd323f40a8c96e70d8. ChatGPT Sites version 13 reports that exact source commit, and production deployment appgdep_6a9a6f2a52d081918f5b700cecc19f78 completed successfully on the public Site. The canonical clean-source release gate passed. Release integrity reported PACKAGE8_RELEASE_PASS with 724 packages and 16 checks; npm audit found zero vulnerabilities. The built-Worker browser, security, responsive, keyboard, and accessibility checks passed. A fresh R13 run in ChatGPT's in-app browser discovered exactly four tools and completed the full public flow: finalized revision-1 observation, proposal creation, visible human approval, guarded apply, revision-2 rehearsal, and all six verification checks. Successful agent mutations refresh the visible page without a manual reload. A fresh apply truthfully reports a revision change; retrying the same idempotency key recovers the same receipt while truthfully reporting no new revision change. R13 preserves the visible-only approval authority, guarded D1 state machine, native dialog, and independent raw-event verifier. R10 retains the historical Chrome trace for the same four-tool contract.

Release and evidence: https://github.com/chvignesh07/focus-contract-studio/releases/tag/webmcp-challenge-2026-r13

What we learned

WebMCP is most valuable when a capability needs the precise state of the live page. Exposing a mutation-shaped function is easy; preserving ordinary authorization, exact review intent, revision safety, accessible fallback, and truthful recovery around that function is the real product work.

We also learned to separate relevance from permission. Precedent can make an agent proposal better and determine whether it is admissible, but it cannot say whether this exact payload was approved. And a renderer cannot prove itself by producing its own expected events—verification needs an independent observation path.

What's next

This contest release proves one dialog family and two visual variants with synthetic precedent. The next evidence step is moderated testing with accessibility and design-system practitioners, measuring task time, correction count, unsupported-proposal rejection, and trust calibration before expanding to more components or organizational roles.

Built with

WebMCP, ChatGPT, ChatGPT Sites, TypeScript, React, Vinext, Cloudflare Workers, D1, Drizzle ORM, Zod, Web Crypto, Vitest, Playwright, Testing Library, and axe-core.

Built With

  • axe-core
  • chatgpt
  • chatgpt-sites
  • cloudflare-d1
  • cloudflare-workers
  • drizzle-orm
  • playwright
  • react
  • testing-library
  • typescript
  • vinext
  • vitest
  • web-crypto
  • webmcp
  • zod
Share this project:

Updates

Submission history