-
-
Live AX Health Score from a real self-scan: seven WebMCP tools discovered and audited across 42 checks.
-
AX Health breakdown across discoverability, clarity, reliability and recovery, and safety.
-
Human × Agent Digital Twin showing 100% capability parity for the scanned WebMCP interface.
-
Failure X-Ray leading into Tool Surgeon, with the primary contract risk and repair workflow.
Inspiration
WebMCP gives websites a native way to expose tools to AI agents, but a valid tool contract is not automatically a usable one. A tool can be hidden, ambiguously named, underspecified, unsafe, or impossible to recover from when it fails. Developers need the agent equivalent of accessibility and reliability testing.
WebMCP Doctor is a read-only diagnostic workbench for that missing layer: Agent Experience (AX).
What it does
Enter any public website URL and WebMCP Doctor loads the real page in an isolated Cloudflare browser session. It discovers the site's WebMCP surface, audits it, and turns the result into one focused report.
- AX Health Score runs 42 deterministic checks across discoverability, tool clarity, reliability and recovery, and safety.
- Agent Failure X-Ray visualizes the path from a goal through discovery, selection, input, execution, and root cause.
- Tool Surgeon diagnoses weak or unsafe contracts, proposes a repaired definition, and compares the same rubric before and after.
- Human × Agent Digital Twin compares actions available in the rendered human interface with actions exposed through WebMCP, highlighting capability gaps.
The scanner never invokes a target site's tools. It inspects definitions and visible controls only.
How we built it
The interface is built with Next.js 16, TypeScript, React, Tailwind CSS, React Flow, Motion, and Lucide. It is deployed as a Cloudflare Worker.
For real-site inspection, the Worker uses Cloudflare Browser Rendering / Browser Run to load the submitted URL and capture multiple WebMCP registration paths:
- imperative
document.modelContext.registerTool(...)calls; - native
document.modelContext.getTools()results when available; - declarative tool forms and their input metadata.
The browser session also extracts visible human capabilities from buttons, links, forms, and controls. A deterministic audit engine then evaluates the collected tool metadata. Score weights are 25% discoverability, 30% clarity, 25% reliability/recovery, and 20% safety.
The Failure X-Ray and parity views are derived from observed metadata; they do not pretend to execute autonomous agents. Tool Surgeon applies explicit repair heuristics, reruns the same checks, and shows the resulting score delta.
The production scanner blocks private and local network targets, blocks unnecessary subresources, and caches results for five minutes. No paid model API, database, login, or analytics service is required.
WebMCP Doctor also exposes its own native WebMCP tools so an agent can operate the diagnostic workflow:
scan_agent_experienceinspect_tooltrace_agent_failurecompare_human_agent_pathssimulate_tool_changerun_ax_testexplain_ax_score
These calls visibly update the same interface a human uses.
Challenges we ran into
The largest challenge was that WebMCP implementations are not uniform. Some sites register tools imperatively during page startup, some expose native inspection, and others use declarative forms. We built a layered collector so the audit remains useful across these patterns without invoking real actions.
Another challenge was keeping the score explainable. Rather than hiding judgment inside an LLM prompt, each finding maps to a deterministic rule with severity, evidence, and a concrete recommendation.
Finally, comparing human and agent capability required careful normalization. Human labels and tool names often describe the same action differently, so the parity engine combines normalized tokens, descriptions, and control context while keeping every match inspectable.
Accomplishments that we're proud of
- It audits real public websites instead of relying on prefilled demo fixtures.
- It makes agent failures visual and traceable instead of returning a generic pass/fail result.
- It can show whether a proposed tool repair actually improves the same test suite.
- The public repository includes 13 unit/security tests, 3 browser tests, and a passing CI quality gate.
- The complete app runs on a free, minimal Cloudflare architecture.
What we learned
Agent-facing interfaces need the same discipline as human-facing interfaces: clear affordances, safe defaults, actionable errors, and recovery paths. WebMCP makes tools discoverable; AX engineering makes them dependable.
We also learned that static analysis becomes much more useful when it is tied to a visual causal model. A developer can understand a red node in the Failure X-Ray faster than a long list of disconnected lint messages.
What's next for WebMCP Doctor
Next steps include reusable CI policies, historical score comparisons, exportable audit reports, deeper declarative-form coverage, and opt-in authenticated scanning for staging environments. The core will remain deterministic, explainable, and safe by default.
Built With
- next.js
- typescript
Log in or sign up for Devpost to join the conversation.