Inspiration
I spent five weeks inside a dispute I could not see into. I sent an evidence file, and never learned whether anyone opened it or which rule the outcome rested on. Losing was survivable.
The blindness was not.
On the other hand, I built that entire evidence file with an AI agent. It read the record and drafted the case, and I could not show a single person what it had opened or why it concluded what it did. My own side of the argument was as unauditable as the decision I was arguing against. The agent was already in the room. Nobody could see it working, including me.
In the 180 days to 2 September 2026, 367 platforms filed 3,804,771,378 statements of reasons to the EU's Digital Services Act transparency database. 43% of those decisions were fully automated, meaning no person looked at them. Every one is somebody sitting where I sat.
Now we send agents into rooms like that one. The companies shipping them settled the question of whether agents act for us. They left the other one open: can you see what yours was allowed to touch, and what it opened. I built the room I wanted to be standing in.
What it does
Two people disagree. Each sends an AI agent to argue for them.
Four agents sit in four separate origins, two advocates and two readers, and a fifth origin holds the record. Each agent holds a different set of tools, and the browser decides which.
To contest a fact, an agent has to open the other side's exhibit and quote it. The page checks that quote against the document and writes a read receipt under that agent's name. Close filing, and the filing tools leave both hands mid-session. Prompting does not bring them back.
confirm ends the case. I never gave it to an agent, in any phase. The agents assemble the record. A named person owns the decision.
Two opposed parties can now put agents into one process and each check what the other's agent was able to do. I looked for a product that already does this: online dispute resolution, agent negotiation, the agent-to-agent protocol drafts. The nearest one runs each side's agent inside its own private tenancy, so no shared surface hands out the capability and neither side can read the other's calls. My search came back empty. That is not proof nothing exists.
How we built it
A disagreement is the sharpest version of the problem WebMCP solves. Two agents work the same record from opposite sides. Neither should reach the other's capabilities, and neither should reach the control that ends the case. A server integration cannot say "this capability belongs to this browser tab's origin and to no other, ever." exposedTo says it, and the browser checks it the moment a cross-origin frame asks.
The record registers each WebMCP tool with exposedTo, naming the one origin allowed to call it. Chrome answers every frame's getTools({ fromOrigins }) request. My application never sees the question.
Each phase owns an AbortController. Closing a phase aborts its signal, and every tool tied to that phase goes with it. An append-only ledger keeps every call that reaches a tool and what came back, refusals included.
A judge can run it three ways: scripted panels with no model at all, a configured model, or their own Claude Code or Codex agent over an MCP-to-WebMCP bridge. The boundary is identical in all three.
Challenges we ran into
WebMCP tool names are unique per document, not per origin. My first registrations collided and left one advocate holding nothing. An empty column looks like a boundary doing its job, which is what makes that kind of bug dangerous. Namespacing every actor's tools fixed it.
Payload limits have to be measured after serialisation. A paged response fit on paper, then broke once JSON escaping doubled every quote inside a claim and cut the result mid-string. The implementation measures the finished string now.
Cross-origin frames and mid-session tool withdrawal both needed a real browser. My unit tests were more permissive than Chrome, which is how the name collision above reached production.
Accomplishments that we're proud of
The Board shows absence. Its manifest lists what each agent holds and what it never received, so a refusal lands in the record instead of disappearing into an error.
An agent cannot dispute an exhibit it never opened, or cite a fact it never assessed. The page checks every quote against its source.
Prompt injection can fool a reader. It cannot conjure a tool the browser never granted, reopen a closed phase, or reach confirmation. The project runs on five deployed origins under 897 automated tests, and in live runs both Claude Code and Codex honoured a mid-phase tool withdrawal without restarting.
Chrome's agent-security guidance names nine agent-side defences. The one mechanism the browser answers for itself, exposedTo, sits somewhere else: the guidance Chrome wrote for site authors. The Board implements the agent-side protections, then puts that browser-held boundary underneath them.
What we learned
The dangerous bugs do not crash. They return a plausible answer built on part of the record, or an empty tool list caused by a name collision. I found both by driving the deployed product and by reviewing it adversarially. Neither showed up in a diff.
Provenance has to be captured while the work happens. A model describing what it remembers doing is not an audit trail. A read receipt written as the call lands is.
What's next for The Board
Two prose fields still clip, so a party can end up reading a cut-off version of why they lost. That one is first. After it: cursor-based access to longer records, more than two readers, and stronger identity around the human confirmation.
Every consequential process that takes an agent's input has to answer two questions. What could it reach, and what did it do. The browser can make both visible, and it is the only party in the room with no stake in the outcome.
Built With
- chrome
- chrome-devtools-protocol
- claude-code
- codex
- deepseek
- javascript
- model-context-protocol
- netlify
- node.js
- pdf.js
- react
- tailwindcss
- typescript
- vite
- vitest
- webmcp
Log in or sign up for Devpost to join the conversation.