Inspiration

Most "AI test agents" still drive the UI the old way: guess a CSS selector, click, hope it doesn't break on the next redesign. When we saw the emerging WebMCP standard — pages exposing their own actions through document.modelContext.registerTool(...), with a proper JSON Schema instead of a DOM tree to guess at — it looked like the missing layer for reliable, AI-native UI testing. If a page tells an agent exactly what it can do and what valid input looks like, testing stops being "click and pray" and becomes "call the tool and check the contract."

What it does

ToolProof is a small Electron "browser" built for one job: point it at any web page and it will test that page through its own tools.

  1. Load — a target URL opens inside an embedded <webview>, just like a normal browser tab.
  2. Discover — it checks the page for a native document.modelContext. If the site doesn't support WebMCP yet, a fallback pass reads the page's own HTML5 form constraints (pattern, required, min/max, type=email, …) and synthesizes an equivalent modelContext on the fly, so any site can be exercised the same way.
  3. Generate scenarios — the discovered tool's JSON Schema is sent to Gemini, which proposes 6–10 realistic test cases: happy path plus boundary and invalid-input cases. If there's no API key, or Gemini fails, a deterministic rule-based generator takes over (missing required fields, out-of-range numbers, regex-violating strings via randexp, invalid enum values) — so the tool never dead-ends.
  4. Run — each scenario is executed live, inside the real page, via document.modelContext.executeTool(...) — not simulated clicks.
  5. Report — pass/fail is compared against the scenario's expected outcome and shown per-scenario plus a summary.

A companion demo page implements a real, native document.modelContext.registerTool flow, used to validate the agent against an actual WebMCP-compliant target instead of only the DOM-inference fallback.

How we built it

  • Electron, with a strict main / preload / renderer split (contextIsolation: true, no nodeIntegration) — the renderer never touches Node directly, everything goes through an IPC bridge exposed in preload.js.
  • The target site itself runs in a <webview>, so the agent's "browser" is a real Chromium instance, not a headless stub.
  • A hand-rolled document.modelContext polyfill is injected into pages without native WebMCP support, so the discovery → generate → run pipeline is identical whether the tools are real or inferred.
  • Gemini (gemini-2.5-flash) turns a tool's inputSchema into concrete test data; a local, dependency-free rule engine (scenario-utils.js) mirrors the same scenario shape as a deterministic fallback.
  • A separate agent-runner (Node + Playwright) explores the same idea headlessly, for CI-style runs outside the Electron shell.

Challenges we ran into

  • IPC serialization: document.modelContext.getTools() can return live objects with circular/non-cloneable references (e.g. a window handle) — Electron's executeJavaScript bridge threw "object could not be cloned" until we normalized every tool down to { name, description, inputSchema } before it crosses the IPC boundary.
  • Timing: tools registered after client-side hydration weren't visible on first check, so discovery needed to poll getTools() for up to ~8s rather than assume readiness on page load.
  • Schema shape drift: some pages return inputSchema as a JSON string rather than an object — discovery normalizes both shapes before generating scenarios.
  • LLM reliability: an LLM won't always return valid JSON or field names that match the schema exactly, so the rule-based generator exists as a hard fallback, not just a nice-to-have.

What we learned

That WebMCP's registerTool / getTools / executeTool contract is a genuinely better testing surface than the DOM — it turns "does the UI still look right" into "does the underlying capability still honor its own contract," which is both more robust to redesigns and closer to what a QA engineer actually cares about.

Built With

Share this project:

Updates

Submission history