Inspiration

Browser agents are powerful, but browser automation becomes brittle the moment a selector changes. A traditional automation script usually just fails, while an autonomous agent may be tempted to guess what to click next.

I built FlowProof to explore a safer middle ground: give browser agents a small, structured QA control plane through WebMCP, preserve evidence when something breaks, and allow recovery only when the current page proves that a safe fallback exists.

The goal is simple: turn a brittle browser failure into a verified regression test.

What it does

FlowProof exposes six native WebMCP tools:

  • create_test
  • start_run
  • get_run_status
  • inspect_failure
  • retry_failed_step
  • export_regression_test

The built-in deterministic /demo-store demonstrates the complete workflow.

A test run intentionally reaches a stale selector:

#checkout-submit

Instead of silently guessing another selector, FlowProof records the failure, captures browser evidence, and inspects the current DOM. Only when the page proves that a stable alternative exists does FlowProof expose:

[data-testid="checkout-submit"]

The failed step can then be retried with that verified selector. After the browser confirms the final Order confirmed state, FlowProof exports the successful recovery as a reusable Playwright TypeScript regression test.

How I built it

The frontend is built with React, TypeScript, and Vite. It registers exactly six tools through the current WebMCP API:

document.modelContext.registerTool(...)

Each tool has a narrow JSON schema and maps to a specific backend operation.

The backend is written in Go. It owns test and run state, target validation, browser orchestration, evidence capture, recovery logic, and regression export.

Real browser execution is handled with chromedp + Chromium. FlowProof captures visible text and screenshots around failures and successful recovery.

The application deliberately does not embed another LLM. The browser agent is the agent; FlowProof provides deterministic execution, evidence, safety boundaries, and recovery primitives.

The production application is packaged with Docker. GitHub Actions builds the container and publishes it to GHCR, and the live service runs on Railway.

WebMCP

FlowProof uses native WebMCP as the agent interface rather than wrapping the application in a custom chat API.

The public production deployment was tested in Google Chrome 151 with WebMCP testing enabled. Chrome detected all six registered FlowProof tools and successfully executed the native workflow against the live application.

This means an agent can create a test, start a browser run, inspect a failure, retry a verified recovery, query the resulting state, and export the regression test entirely through structured WebMCP tools.

Challenges

One of the biggest challenges was making the demo genuinely deterministic while still using a real browser.

The stale selector needed to fail for real, browser evidence needed to be preserved, and recovery had to be based on observed DOM state rather than a hard-coded success shortcut.

I also encountered a WebMCP compatibility issue in Chrome where the execution callback could be invoked without a context object. The original implementation assumed the context always existed. I added compatibility handling and regression coverage while preserving AbortSignal forwarding when the browser provides it.

Deployment had its own challenge: the original cloud build path failed before reaching the Docker build. I moved container building to GitHub Actions and GHCR, then deployed the published image directly to Railway.

What I learned

The most important lesson was that exposing a browser action as a tool is not enough. Agent-facing browser tools need strong semantics around failure, evidence, state, and recovery.

WebMCP provides a clean browser-native boundary for those tools. Keeping each tool narrow also makes the workflow easier for both agents and humans to reason about.

I also learned that recovery is much more trustworthy when the system can explain not only what selector it wants to use next, but why that selector is allowed based on current browser evidence.

What's next

FlowProof currently focuses on a deterministic checkout fixture so the entire recovery path can be evaluated reliably.

The same architecture could be extended to larger QA suites, richer evidence comparison, multiple browser environments, CI regression execution, and broader application-specific WebMCP test contracts.

The core principle would remain the same:

evidence first, recovery second, reusable regression test last.

Built With

Share this project:

Updates

Submission history