Inspiration

What it does

How we built it

Challenges we ran into

Accomplishments that we're proud of## Inspiration

A WebMCP tool publishes an input schema. That schema is a promise to the agent: these are the fields, this one is required, this one has to be one of three values.

Nothing enforces it. Not the browser, not the API. That turns out to be deliberate, and the spec's own issue 92 has been debating whether it should change since February.

I found it by accident while testing a tool I was writing. I sent it a quantity of "two" as a string where the schema declared a number, and it went straight through to my handler. Then I sent negative 999 cups against a declared minimum of 1. That went through too. The schema I had carefully written was decoration.

So I stopped building what I was building and went to measure how far it goes.

What it does

lapse measures the gap between what a WebMCP tool promises and what it actually accepts.

The targets are not toys written to fail. They are the official Chrome WebMCP reference demos from GoogleChromeLabs, vendored unmodified under Apache-2.0. The coffee shop, the pizza maker, the flight search, the ticket booking, the real estate map. The code Chrome ships to teach people how to build with WebMCP.

lapse enumerates every tool through getTools, reads each declared schema, and generates input that schema forbids. Nulls where a string was declared. Booleans and arrays and objects where a scalar was declared. Values outside a declared enum. Numbers past a declared maximum. Required fields removed. A ten thousand character string. Control and bidirectional characters. Text written to sound like an instruction aimed at the agent. Then it calls the tool and records what happened.

299 probes across 23 tools on 8 sites. 225 violations. 17 of 17 imperative tools accept input their own published schema forbids, at 78 percent of every scored probe.

The clearest single case: the pizza maker declares topping as a fixed list. Send it a value that is not on the list and it replies that it added one off-menu topping.

The violations cluster identically across every site. That is what a platform property looks like, not seven separate authoring mistakes.

Why this is a strong fit for WebMCP

It could not exist anywhere else. It runs inside the browser, against live tools, using the API's own enumeration and execution surface, and the thing it measures is a property of WebMCP itself that no existing tool reports. Chrome ships an evals CLI that tests whether a model picks the right tool. Nothing tests whether a tool honours what it published.

How it makes for a better experience

A developer who wants to know whether their tools are safe under an agent has one option today, which is to write eval cases by hand and read the output. lapse gives a verdict in under a minute, per tool, per constraint, with the failing input printed next to what the tool did with it. You do not have to guess which of your fields is unguarded. It tells you.

What people and agents can do together here

lapse registers four WebMCP tools of its own, so the audit is driven by conversation rather than by clicking. Ask which sites are loaded and what each one exposes. Ask it to run the audit. Ask for every violation on one site. Ask it to export the report for a bug filing.

Auditing a site used to mean a person writing test cases one at a time and reading them one at a time. Here the agent generates and executes hundreds against live tools while the human watches the result form on a shared page. Neither side could do it alone at that speed. The agent has no way to reach into a page and call its tools without WebMCP, and the person has no patience for 299 hand-written cases.

How I implemented WebMCP

lapse uses the API in both directions.

As a client it calls getTools to enumerate tools across a same-origin frame tree, executeTool to drive them with generated input, toolchange to track registration, AbortController to cancel calls that stop responding, and it reads annotations to check whether a tool that echoes agent-directed text sets untrustedContentHint.

As a server it registers its own:

document.modelContext.registerTool({
  name: "run_audit",
  description: "Run the full contract audit across every vendored demo. Returns probe, violation and not-measured counts.",
  inputSchema: { type: "object", properties: {} },
  execute: async () => (await runAudit()) || "audit produced no findings",
  annotations: { readOnlyHint: false }
});

Four tools in total: list_targets, run_audit, get_findings, export_report.

Challenges, and the verdict that saved me twice

Every probe lands as HONOURED, VIOLATED, or NOT MEASURED. A probe that could not be generated, could not be run, or never returned is recorded as unmeasured. It is never counted as a pass.

That third verdict caught two false findings that would otherwise have shipped.

The first version read inputSchema as an object. It actually arrives as a JSON string. Every schema-derived probe silently generated nothing while the tool confidently reported five violations as though that were the answer. The arithmetic gave it away: four tools, nine probes, exactly the two unconditional checks each.

The second was worse because it flattered me. Declarative tools appeared to validate perfectly, which would have been a lovely headline. So I suppressed the form navigation that was cutting those tools short, reran, and the outcomes changed. The guard was altering the thing it was supposed to leave alone. I removed it and marked six declarative tools unmeasured rather than publish the better number.

What I learned

This is not a bug report. The working group already knows the schema is advisory. Issue 92 on the spec repo, "Who owns the validation layer?", has been open since February and argues that inputSchema is only a semantic hint, that validation falls to each tool author individually, and that the browser should probably own it instead.

What that discussion has never had is a measurement. It has been argued from experience and intuition for six months. lapse gives it a number: across the reference demos Chrome itself ships to teach the API, 17 of 17 imperative tools have unguarded input paths, and 78 percent of probes that violate a declared constraint are accepted anyway.

The separate question I could not settle is the declarative side. Calling a declarative tool submits its form and navigates the document, so it cannot be probed twice. I tried suppressing that navigation, watched it change the outcomes it was meant to leave alone, removed it, and marked those six tools unmeasured. Whether declarative validation is genuinely stronger remains an open question, and lapse says so rather than guessing.

What's next

More targets, including the demos that need a build step. Chrome publishes an eval JSON format for its CLI, and lapse should read and write it so the same cases can run in a browser against live tools instead of only in a terminal.

What we learned

What's next for lapse — what WebMCP tools actually accept

Built With

Share this project:

Updates

Submission history