Inspiration
Every judge on this panel works at a company that lives in an issue tracker. "Triage my backlog, dedupe these, split this ticket, reprioritize the sprint" is a request any engineer has typed into Slack at 5pm on a Friday. It's also genuinely tedious for a person and genuinely tractable for an agent, but only if the agent can act with real authority inside the tool instead of clicking through a UI it's guessing at from a screenshot.
We also noticed a gap. The official showcase and the community demo list are full of storefronts, flight bookers, and games. Nothing engineering-tool shaped. Nothing where the agent has to hold a real permission boundary, share an undo stack with a human, or watch its own available capabilities change as the page changes under it. That felt like the actual hard part of WebMCP, the underexplored part, so we built toward it instead of around it.
What it does
Cadence is an issue tracker where an agent is a real teammate on the board, not a chatbot bolted onto the side of one. Ask it to triage an untriaged backlog, find and merge duplicate bug reports, split an oversized epic into sub-issues, or reprioritize a sprint, and it does real analysis against real board state through the same functions the UI itself calls.
Here's the part that's actually WebMCP-specific, not just "an app with an AI feature": the registered tool set changes live with what's on screen. Select an issue and add_comment, split_issue, and link_issues register. Apply a filter or multi-select and bulk_update shows up. Switch to cycle view and summarize_cycle appears. None of that is expressible as a static, pre-declared tool list, because a server-side MCP server has no idea what a human currently has selected in their browser. It can't.
Agents also carry real identity and real permission scoping. A workspace owner grants each agent a scope, read, triage, write, or full, and that grant filters which tools registerTool even offers before the agent ever sees them, live. Every agent edit lands on the same undo stack the human uses too. Cmd+Z undoes an agent's change exactly like it was your own, because both write through one shared reduce() function against one shared state.
How we built it
React, TypeScript, and Vite on the client. Cloudflare Workers plus a Durable Object (SQLite-backed, free-tier eligible) as the authoritative backend, syncing over the WebSocket Hibernation API so multiple humans and agents can work the same board live.
webmcp-kit, a small library we published alongside this entry: defineTool/registerTools with full JSON-Schema-to-TS inference, and useScopedTools, a React hook that registers a tool set while a condition holds and cleanly unregisters when it stops. That hook is what makes Cadence's dynamic tool surface a one-liner instead of a pile of manual AbortController bookkeeping.
Every mutation is a pure (state, actor, action) → state function in one reducer, shared by the client store, the Durable Object, and every WebMCP tool handler. There's exactly one place a board mutation is defined, which is also what makes the shared human/agent undo stack possible. Only one code path to undo.
Seed data is deliberately messy on purpose: real duplicate bug reports worded three different ways, vague untriaged titles, missing estimates, so the higher-order tools have genuine work to do instead of a tidy board with nothing left to fix.
A hash-chained audit log on the Durable Object records every mutation with actor attribution, queryable via get_activity.
Challenges we ran into
ChatGPT's in-app browser only supports a subset of the spec: imperative registration only, no declarative HTML tool attributes, and no cross-frame tool discovery even for same-origin iframes. We designed single-origin from day one because of this, and honestly it shaped the whole architecture more than anything else did.
Then there was the embarrassing one. Our seed data computed Date.now() at Cloudflare Workers module scope instead of inside a request handler. Turns out Workers doesn't guarantee wall-clock time outside a request's I/O context, so every "recent" issue and cycle date came out silently wrong, offsets from epoch, not from actual "now". We only caught it by inspecting live state over a raw WebSocket and noticing the numbers just didn't add up.
find_duplicates had a similar quiet failure. It used Jaccard similarity on tokenized title and body text with a default threshold that, on inspection, was too strict to ever catch our own deliberately seeded duplicate cluster. Real paraphrased duplicates scored around 0.23 to 0.28 against a 0.4 cutoff. Tuned it down and documented why right in the tool description, since it's a real limitation an agent calling the tool should actually know about.
And demo-state entropy: a Durable Object's seed loads exactly once, ever, so every round of testing permanently mutated the same live board. Later testers kept hitting an already-triaged, already-deduped board with nothing left to demonstrate. Fixed it by giving every visitor their own isolated board, auto-fresh on first visit, plus a visible in-app reset button, instead of trying to guess when it's "safe" to reset something shared.
Accomplishments we're proud of
A genuinely dynamic tool surface, verifiable live in Chrome's Model Context Tool Inspector: select an issue and watch the registered tool count change in real time, no faith required.
Per-agent permission scoping that changes what tools even exist for an agent, not just what it's allowed to call, enforced before registration ever happens.
A free, LLM-free benchmark comparing tool calls against real Playwright-driven DOM interaction on the live app. Single actions cost about the same either way, but a 7-item bulk task runs roughly 3x faster over tool calls, because the DOM path re-pays a fixed per-item navigation cost that a tool call just doesn't.
Real accessibility work, not just a claim. Full keyboard operability, and a code-level WCAG audit that actually found and fixed two real bugs, custom buttons not activating on Space, confirmation dialogs with no Escape-to-dismiss, instead of asserting AA compliance from a README and hoping nobody checks.
What we learned
The most interesting property of WebMCP isn't that it lets an agent call functions instead of scraping a page. It's that tool registration lives in the same runtime as the UI state, so the tool surface can be a live function of what's actually on screen, and permission boundaries can be enforced at the moment of registration rather than only at call time. Neither of those is something a server-side MCP server wrapping a REST API can do, no matter how well-designed that API is.
What's next
Publishing webmcp-kit to npm proper (it's a GitHub dependency right now), a full screen-reader session to go with the code-level accessibility audit, and pushing the duplicate-detection scoring past bag-of-words Jaccard now that we know exactly where its ceiling is.
Built With
- cloudflare-durable-objects
- cloudflare-workers
- css3
- html5
- model-context-protocol
- node.js
- react
- typescript
- vite
- webmcp
- websockets
Log in or sign up for Devpost to join the conversation.