Why this use case is a strong fit for WebMCP

Identity-mapping rules often exist only as an unsaved draft in an administrator's browser. The failures that matter tend to appear at the edges: a case-sensitive check puts a contractor in the employees group, a missing manager becomes an empty string, or stale directory data overrides HR.

At that point, the page is the only place that knows the current draft, its revision, its fields, and its rules. WebMCP lets the page expose that information through five structured, least-privilege tools. The agent can work with the page's actual semantics instead of guessing from labels or DOM layout, even though the draft has not reached a server.

How it creates a better user experience

The administrator continues editing the draft in the normal page. The agent reads that exact draft and its revision, then stages three safety rules. Those rules stay pending until the administrator reviews their cards and clicks Confirm all.

Once confirmed, the agent finds the smallest set of synthetic users that exposes every violation. For each result, the page shows which source value won, which value lost, and why. The agent can preview a redacted fix, but the administrator makes the edit.

An edit immediately marks any dependent evidence STALE. The agent must test the new revision before it can prepare a GREEN review packet. There is no Save or Apply tool; applying the mapping remains a normal page control outside WebMCP.

What people and agents can do together that was difficult or impossible before

The administrator and agent can reason about a draft that exists only in the browser, using evidence that expires when its basis changes. Expressions, source priority, and confirmed rules can each invalidate an earlier result. Every evidence record therefore carries its revision and dependencies. When it becomes stale, the page strikes through the old rows and blocks the old packet.

The agent handles the exhaustive search for edge cases across the synthetic population. The administrator confirms the rules and makes the actual edit. That division prevents proof from an older draft from being presented as proof of the current one.

How we implemented WebMCP

The top-level document registers five tools with document.modelContext.registerTool({ name, description, inputSchema, execute }) in app.js. They read the session and draft, stage rules, find a minimal counterexample set, preview a patch, and prepare the review packet.

Strict schemas reject malformed input and extra fields. Failed calls leave the committed draft unchanged. Responses have size limits and redact CANARY_ fixture strings.

We initially tried to wait for human approval inside one long-running tool call. In our ChatGPT in-app-browser testing, pending calls ended after roughly 22 seconds. The implemented flow therefore stages the rule cards and returns. After the administrator responds, the agent reads the new state on its next turn.

The suite has 310 passing tests. On Chrome 152, we also run fresh-session registration checks and a 12-round trace that combines real DOM edits with WebMCP calls.

Inspiration

Identity teams merge records from Active Directory, an HR system, and an identity provider such as Okta before the resulting profiles determine access. The merge rules may still be an unsaved browser draft while the administrator works on them.

Small mistakes can have large effects. A case-sensitive condition may classify a contractor as an employee. A missing manager may turn into an empty string instead of null. Stale directory data may take precedence over HR. These problems are easy to overlook without the right edge cases.

An AI agent can help search for those cases, but unrestricted UI control introduces a different set of risks. The agent may guess what a field means, use evidence from an older draft, or act before a person has reviewed the result.

What it does

Two boundaries are important up front. Unsaved-draft preview already exists as a first-party product pattern. A different page-local agent with access to the same state and rules could also run this engine. IdentityMap Witness contributes a page-authored safety contract and an evidence lifecycle that causes old proof to expire.

IdentityMap Witness reviews an identity-mapping draft before it is saved. A witness is the smallest set of synthetic users required to expose every safety rule that the current draft violates.

The workflow is:

  1. The agent reads the exact unsaved draft and its revision.

  2. It stages three safety rules. They remain pending until a person reviews the cards and clicks Confirm all.

  3. The agent finds the fewest synthetic users that demonstrate every violation. The result explains which source value won, which lost, and why.

  4. The agent previews a redacted fix, and the person performs the actual edit in the page.

  5. The edit marks dependent evidence STALE. The new revision must be tested before the agent can prepare a GREEN review packet.

The agent never receives a Save or Apply tool. Apply mapping (manual page control) remains outside the WebMCP surface and is not used in the demo.

Why WebMCP

The page is the source of truth for the live draft. It owns the fields and rules, tracks the revision, and decides which operations are allowed. WebMCP exposes that knowledge through five structured, least-privilege tools instead of making the agent infer it from labels or DOM structure.

The tools return redacted evidence tied to a particular revision and reject calls made against old state. This allows the administrator and agent to work with data that has not reached a server while confirmation and editing remain visibly under human control.

A general browser agent could inspect the interface. The difference here is that the page itself defines the contract, which makes the workflow reviewable and prevents stale conclusions from being silently reused.

How we built it

The application uses vanilla JavaScript, ES modules, and Node 21+. Five document.modelContext tools read the session, stage rules, find a minimal counterexample set, preview a patch, and prepare the review packet. The browser, test suite, and evaluation all use the same deterministic engine.

Tool inputs use strict schemas, so extra or malformed fields are rejected. A failed call does not change the committed draft. Responses are size-limited and redact CANARY_ fixture strings. Before-and-after diffs are minimized for firstName, lastName, email, displayName, and managerId. This is a synthetic tripwire, not a general-purpose PII detector. Evidence records include the revision and dependencies used to determine whether they remain valid.

There are 310 passing tests. Chrome 152 runs three fresh-session registration checks and a 12-round trace with real DOM edits and WebMCP calls. A hand-audited oracle verifies that the first search returns {P2, P3, P4}. This is one of two audited minimal sets of size three, and its four violation rows cover all four seeded defect classes.

Challenges we ran into

The central problem was that a correct result could become unsafe after the draft changed. An edit to an expression, source priority, or confirmed rule may invalidate earlier reasoning. Evidence is therefore tied to those exact dependencies; relevant edits strike through the old rows and prevent the old packet from being used.

Human approval could not remain pending inside one tool call. In our own ChatGPT in-app-browser test on 2026-08-29, pending calls ended after roughly 22 seconds. That observation is not part of the repository's automated evidence. The current tool stages the cards and returns, the administrator confirms them, and the agent reads the updated state on its next turn.

We also turned stale buttons, malformed inputs, hostile text, and empty source values into executable failure cases.

What's next

The demo currently uses eight synthetic personas, tab-local state, and an exhaustive search sized for this fixture. It is not connected to a real identity provider or a save operation. Unsaved-draft preview is not new on its own, and another page-local agent could run the same engine. The contribution here is the page-authored safety contract and the lifecycle that expires old proof.

The Browser-Use and Full-CDP comparison arms have been designed but not run, so we make no comparative claim.

Next, we would connect a real provider behind the same redaction boundary, test larger datasets with consent, persist review history, and add signed audit receipts where cross-session trust is required.

Try it

  1. Open https://identitymap-witness.onrender.com in ChatGPT's in-app browser and enable site tools for the page.

  2. Pass the visible consent gate. Select Copy prompt 1 — setup, paste the copied prompt into chat, and send it.

Built With

Share this project:

Updates

Submission history