ChenChess Coaching Board: Grounding a WebMCP Agent on a Live Chessboard

Inspiration

A player studying a game moves pieces with the mouse and asks questions by talking to a model. Those are two separate channels, and the language channel is full of pointers the model can't resolve on its own — why is this bad, what about that instead, is the line I tried still winning — all referring to board state the model never actually saw. WebMCP closes that gap: the page itself registers tools with document.modelContext, so an agent like ChatGPT or Claude can call them against the live page and the player's own session, instead of guessing from a transcript.

The interesting problem wasn't wiring up tool calls. It was making sure an agent with real tools couldn't say anything about the board that wasn't true — the same grounding problem the rest of the coach has, now with a model that can also act.

The coach-engine logic underneath

Every WebMCP tool call ultimately reads from one deterministic pipeline, not from the model's own reasoning:

  1. Each position in a game is analyzed by Stockfish (objective best move and evaluation) and Maia-2 (what a human at a given rating would likely play) in parallel.
  2. A pure Rule Extractor compares the two and picks Critical Moments — plies where the position's win probability swung enough to matter. Evaluations are converted from centipawns to a win-probability curve first, since a swing means different things near equal and near winning:

$$ \text{WinProb}(cp) = \frac{1}{1 + 10^{-cp/400}}, \qquad \left|\,\text{WinProb}(cp_{i+1}) - \text{WinProb}(cp_i)\,\right| \ge \tau $$

  1. Each Critical Moment becomes a typed fact record (best line, refutation, verdict) before any prose exists, and a validator gate rejects any later explanation that doesn't match those facts ply-for-ply.

The WebMCP layer doesn't get to skip this. Every board tool call returns a Coaching Board Snapshot — the same typed facts, plus which line is active and where the viewed ply sits on the game's own line — so the agent is handed grounded state on every call instead of being trusted to ask for it.

How I built WebMCP support

  • One authored map, not a second one. Coach tools already carried a coachToolSurface map of which tools the installed Coach App vs. the model may call. Adding the Coaching Board meant widening that map with a "web" target rather than maintaining a parallel list — the two surfaces share list_critical_moments, open_review_moment, and evaluate_player_line against the same data, distinguished by per-tool description, not by renaming.
  • Register inside the auth gate, never at module load. Board routes pass through an asynchronous Beta Access check before a player is known. Tools register on an effect keyed by the authorized player ID and tear down through the AbortSignal registerTool accepts, so there's no window where an agent can discover and call board tools before identity resolves or after sign-out.
  • Push the grounding policy into every result, not just instructions. Descriptions carry the rules once, at registration; each result also carries a constraint block for the facts it just returned, because nothing in WebMCP obliges an agent to re-read state before answering — an agent that skips the read produces fluent, confident, wrong coaching a player can't detect.
  • Keep durable writes off the surface. The agent can stage a game import; only the player's own button press commits it.

Challenges I faced

  • Proving it actually grounds, not just that it compiles. A scripted deixis suite driving the board through a real WebMCP host measured 7/8 on referring expressions ("that line," "the other one") but over-called tools on a non-deictic trap prompt roughly one run in five — enough signal to trust the mechanism, not enough to gate a release on yet.
  • No production host to test against on demand. Storybook turned out to be a legitimate WebMCP host for driving the real board tools headlessly through Playwright with --enable-features=WebMCP, since neither ChatGPT's desktop app nor a locked devtools profile could be automated directly.
  • A stale tab looks identical to a broken one. A board left open for hours answered every tool call with a bare "unavailable," which looked like a server-side bug; it was a client holding a dead connection. Fixing it meant turning any client-side throw into a typed unavailable result (rather than an opaque failure) so the agent — and future debugging — could tell "stale" from "actually broken."
  • A withdrawn move can wedge every future write. A Slot Marker left behind by a since-withdrawn line kept blocking every Review Moment the board tried to write afterward; the fix surfaces the board's refusals in a banner instead of failing silently on the next unrelated call.

The throughline: giving a model real hands on the board is only safe once every fact it can act on, and every fact it hands back, traces to the same deterministic pipeline the rest of the coach already trusts.

Built With

Share this project:

Updates

Submission history