Inspiration
This was my first project using WebMCP, and playing around with the ChatGPT client and examples in the OpenAI showcase, one question kept coming to mind: What if a tool doesn’t behave the way it claims to?
A tool might describe itself as read-only while still changing browser state, writing to storage, or sending a request that modifies data. Developers need more than a list of available tools. They need a way to verify what those tools actually do.
I wanted to make that comparison visible, repeatable, and useful to both humans and agents.
WebMCP is a strong fit for this problem because it gives agents a structured way to discover and invoke capabilities directly from a website. ToolTruth uses that same interface as the subject of inspection, evaluating not only what a WebMCP tool claims to do, but what actually happens when it runs.
What it does
ToolTruth inspects a WebMCP-enabled website inside an isolated browser session.
After the user enters a URL, ToolTruth:
- Discovers the WebMCP tools exposed by the website.
- Reads each tool’s description, input schema, and annotations.
- Generates safe synthetic input for the selected tool.
- Invokes the tool inside a disposable browser.
- Captures output and observable side effects.
- Compares the browser’s state before and after execution.
- Produces a verification result with supporting evidence and a suggested repair.
The evidence workbench includes a timeline, browser-state diff, network activity, and runtime logs.
ToolTruth also exposes its own WebMCP tools so an external agent can inspect the current run, verify one tool, run every verification, and focus the interface on specific evidence while the human follows along. ToolTruth can also be used by developers to test their tools and make sure that they work as intended.
This creates a better user experience by turning tool execution into a transparent, guided workflow. Instead of trusting a description or sorting through raw hard-to-understand logs, users receive an clear verdict, the evidence behind it, and a suggested next step. It also enables a new kind of collaboration between people and agents. An agent can handle repetitive tasks such as discovering tools, running verifications, and locating relevant evidence, while the human reviews the consequences, applies context, and decides whether the tool should be trusted. Reproducing that process manually would require developers to inspect several kinds of browser activity separately for every tool.
How I built it
I implemented WebMCP in two directions. ToolTruth consumes the WebMCP tools exposed by inspected websites so it can discover and test them. It exposes its own WebMCP tools so an external agent can operate the verification workflow alongside the human user.
I built ToolTruth with Next.js, React, TypeScript, Tailwind CSS, and shadcn/ui components.
Stagehand provides the browser automation layer used to discover and invoke WebMCP tools. Inspections run on either disposable local browsers or on Browserbase's browsers in the cloud. OpenRouter and Vercel AI SDK are used for simple implementation of AI features and access to different models without having to make drastic changes to the code.
For every verification, ToolTruth takes a baseline snapshot, calls the selected tool, and then captures the resulting state. It compares signals including page content, rendered pixels, form fields, cookies, local and session storage, network requests, console messages, and runtime errors.
A deterministic analysis layer performs some checks to catch common errors and red flags. When an OpenRouter model is configured, AI helps generate harmless schema-compatible inputs and turns the collected evidence into a concise explanation and repair suggestion. The model is never allowed to override the deterministic verdict, and ToolTruth can fall back to schema-derived inputs and rule-based summaries without AI.
The website streams discovery and verification progress to the interface so users can see what the browser is doing as evidence is collected.
Challenges I ran into
One of the hardest challenges was defining “observable behavior.” Once a tool is called, it can have effects that extend beyond return data. They can appear as changes to the DOM, navigation, storage writes, cookies, network mutations, console errors, or visual changes. Capturing these signals without overwhelming the user or bloating the UI required careful thought and design.
Another major challenge I ran into was dealing with the Browserbase Free Plan limits. You only get 60 minutes of browser time on the free plan and Browserbase will charge a minimum of 1 minute per session no matter the session length. To work around this, I had to create and use multiple accounts and manage different API keys for both local testing and production.
Security was another major challenge. ToolTruth opens URLs and runs third-party tools, so I had to treat the inspected website, its metadata, and its output as untrusted. I added URL validation, restrictions against private and metadata addresses, network controls, secret redaction, and hashing for sensitive browser values.
Browser lifecycle management also proved more complicated than expected. Discovery and verification are asynchronous and stateful, so I had to handle cancellation, disconnected event streams, stale results, concurrent operations, timeouts, and reliable cleanup of disposable sessions.
Finally, I had to design one workflow for two audiences: a human using the visual workbench and an agent operating ToolTruth through WebMCP. Both needed to share the same state and receive consistent results.
Accomplishments that I'm proud of
I am proud that ToolTruth completes the full journey from discovering an unfamiliar WebMCP tool to producing an evidence-backed verification result.
As a freshman who just started high school, it was difficult for me to find time to work on this application. I spent time during the day and stayed up late at night to finish this project.
Rather than showing only a simple pass / fail status, it lets users & agents inspect the underlying timeline, state changes, network requests, and logs. This makes the verdict understandable and gives developers something concrete to debug.
I am also proud that ToolTruth does not merely inspect agent tools, but rather exposes a semantic tool surface that lets another agent operate the inspection workflow while keeping the human in the loop.
Most importantly, safety was built into the architecture. Inspections are isolated, sensitive values are protected, untrusted content is not treated as instructions, and AI-generated explanations cannot change the deterministic result.
What I learned
Participating in this hackathon taught me about WebMCP, what it is, how it works, and how to use it. I had a lot of fun playing around with the ChatGPT client and asking it to invoke tools in WebMCP supported webpages.
Additionally, this was my first project where I used an agentic development process with Codex. Usually I would write most of the code by hand. However this time, Codex would generate the code and I provided the ideas and reviews. I also kept a Notion page with all the important information I wanted it to know so that it would never drift. It was quite different compared to how I usually build but I learned a lot about the process and it helped me iterate extremely fast.
I also learned that no single signal tells the whole story. A tool may leave the visible page unchanged while executing unintended side effects, so meaningful verification requires several sources of clear evidence and proper observation.
AI was most effective when I gave it a constrained role. For example, generating safe test data and explaining evidence. Deterministic rules are most effective when used for catching obvious red-flags. AI responses were used for more difficult cases that required more reasoning and context analysis.
What's next for ToolTruth
To make the application both easy to use and build, I decided to skip implementing authentication or syncing the information to a database. However, these features are definitely on the roadmap.
Some other features / improvements I have in mind:
- Improving accuracy and edge cases
- Improve handling for tools that require certain data to be set before they can used
- Validating outputs against declared schemas
- Testing additional annotations and consent requirements
- Detecting unexpected destination domains
- Allowing developers to define custom behavioral policies
- Adding automatic syncing so that tests rerun when a tool changes
In the long run, I hope that ToolTruth can become the shared trust layer for the agentic web where agents discover and test capabilities, developers diagnose issues, and humans review evidence before deciding which tools deserve access to meaningful actions.
Built With
- browserbase
- nextjs
- openrouter
- react
- sse
- stagehand
- webmcp
Log in or sign up for Devpost to join the conversation.