The problem

Rostering is a constraint problem that looks like a conversation. A manager says "give Tom Friday off" and means "re-solve a 210-variable integer program under thirty-three rules, and tell me if it can't be done." Language models are very good at the conversation and very bad at the program. Ask one to build a ten-person week with coverage minimums, keyholder cover, eleven-hour rest, contract ceilings and floors, and five separate absences, and it will produce something that looks like a roster. It will not be optimal, it will usually break a rule, and when the week is genuinely impossible it will hand you a schedule anyway, because producing plausible text is what it does. Nobody notices a wrong roster in the tool. They notice it on a Friday night when the cafe cannot legally open.

What we built

A web page that gives the agent the thing it does not have: a real mixed-integer solver, HiGHS, compiled to WebAssembly and running in the tab. The agent talks to the manager and edits the rules; the solver decides. Every answer is one of exactly three: optimal, meaning no better roster exists under these rules; infeasible, meaning none exists at all; or timeout. There is no fourth branch where it guesses.

Why this is a strong fit for WebMCP

Three properties of this problem are the argument for putting tools in the page rather than behind an API. First, the engine has to be where the data is: rosters contain names, pay rates and the reasons people are absent, which is the least uploadable data a small business owns, and WebMCP lets the computation happen inside the tab the manager already trusts. There is no server in this project at all - no account, no upload, no database - and it ships as a static export, so "the roster never leaves your browser" is true of the deployment as well as of the code. Second, the answer has to be a fact rather than a message: an in-page tool returns a value the model did not author, and "infeasible" plus six rule ids is not something an agent can talk its way around. Third, both parties need to be looking at the same thing - the manager watches the grid fill in as the agent solves, and the agent sees the rules the manager just changed.

How it creates a better user experience

Granting one person a Friday off in the seeded week produces this, in about 200 milliseconds and 25 solver probes: "These rules cannot all hold at once: a keyholder must open, a keyholder must close, eleven hours rest between shifts, S6 cannot work Fridays, S9 is away Thursday and Friday, and S2 asked for Friday off." Six rules out of thirty-three. Four other people are also away that week and are not mentioned, because they are not part of the clash - and proving that is most of the work. The page then says which of the six would be enough to relax on their own: five would, and the rest rule would not, because a second rule independently forbids the same double shift. That distinction matters more than it sounds: without it the page would send a manager off to relax a rule that leaves them exactly where they started. Then it stops. Choosing which rule gives way is a judgement about people, and neither the agent nor the solver is entitled to make it.

What people and agents can do together that was difficult or impossible before

A manager can describe a change in plain language and get back a proven-optimal week, or a proof that no week exists, instead of hand-checking a roster a model invented. When it fails they get six named rules and what relaxing each one would buy. A member of staff can ask "can I have Thursday off?" and get an immediate verdict with the blocking rules named, rather than waiting on a manager. They can ask "will anyone take my Friday night?" and get the list of colleagues the solver has verified can take it without breaking the week - each one checked by re-solving the entire roster with that person pinned to that shift - rather than asking in a group chat. And all of it happens over data the agent is never allowed to see.

How we implemented WebMCP

Every capability is declared exactly once, in a single action registry. The React UI renders its controls from it and the WebMCP layer registers its tools from it, so there is no second code path for agents and the two surfaces cannot drift apart - the standing objection to WebMCP, that a site ends up maintaining two versions of itself until the agent is calling a tool that lies. A test fails the build on any capability that is neither exposed to agents nor carries a written reason for being withheld; three are withheld and each says why in the source.

The tool surface is a function of state rather than a constant. Tools are registered when they become possible and aborted when they stop being: explain_conflict only exists while the roster is actually infeasible, publish_roster only once there is a clean solved week, and the manager and staff views expose different surfaces entirely. The browser's site-tools count changes as you work, because the tool list is the state of the page.

WebMCP has no confirmation API - the spec lists user prompting as an open question, and requestUserInteraction, which several write-ups describe, does not exist in any shipping build. So publish_roster returns a promise the page resolves on a real click: the agent's call stays open, the card shows exactly what will change, and declining returns declined with nothing done.

Everything an agent receives passes through a redaction boundary. It sees S3, roles, skills, counts and structure; never names, private notes, absence reasons or pay. That is tested as a property rather than asserted as a policy: every private field is replaced with a unique random token, every tool is called many times with sensible and hostile arguments across both roles and both feasible and infeasible states, and one token appearing in one result fails the build. A negative control plants a leak on purpose and requires the same check to catch it.

Two things we could only learn by running it. Chrome 151 calls execute with one argument, so the AbortSignal the spec describes is not there and the binding supplies its own. And the spec says the absence of readOnlyHint is itself the write signal - true of the spec, false of the shipping consumer, because ChatGPT's browser reads the property to build its read/write split, so omitting it on writes displayed our six-tool surface as "3 read, 0 write tools" and the writes vanished from the affordance a person uses to decide whether to trust the page.

Proof

190 unit and integration tests, 34 assertions driven against real Chrome with WebMCP enabled, and 55 eval steps in Chrome's own expectedCall format that run with no model and no API key. Every solve returns a receipt - a hash of the canonical model, the solver version, the status and a hash of the schedule - and a committed script re-solves and compares them, so determinism is enforced rather than asserted. Everything the project claims is listed in CLAIMS.md with its evidence tier and the command that re-derives it, including an explicit list of things we are not claiming.

Built With

  • highs
  • indexeddb
  • next.js
  • puppeteer
  • react
  • tailwindcss
  • typescript
  • vercel
  • vitest
  • web-workers
  • webassembly
  • webmcp
Share this project:

Updates

Submission history