-
-
Next week's rota as you find it: 73% covered, fifteen unfilled slots, and one rule breach that nobody has spotted.
-
The agent calls the page's own rule engine and rings the two shifts it means: nine hours' rest where eleven are required.
-
Fourteen proposed assignments render as dashed cards on the shifts they affect. Nothing has been applied to the rota yet.
-
The consent surface. Every edit is a checkbox, with before and after coverage, breaches and cost, plus what approval would leave unfixed.
-
Every registered tool with its JSON Schema, runnable by hand. The badge shows whether WebMCP is native here or the bundled polyfill.
-
Every tool call recorded: arguments, result, timing, whether it was read-only, the edits it staged, and who called it. Exportable as JSON.
-
Hours against contract, and who is carrying the weekends and the closes. Managers get this wrong by hand, and staff notice.
-
Wage cost by day and by person against a GBP 5,000 weekly budget, with overtime priced at 1.5x. The agent reads this before proposing.
-
The same week in dark appearance. Light and dark are both first-class, and the theme follows your system by default.
Inspiration
I didn't want to build another "AI does your admin" demo. They're easy to make and they all feel the same, because the interesting part gets skipped. An agent doing a job for you is only half a product. The other half is how you check it.
Shift rotas turned out to be a good place to ask that question. A rota isn't really a spreadsheet, it's a promise to fourteen people about their week. Get it wrong and someone's childcare falls through. So it's exactly the kind of work you'd happily hand to a machine, and exactly the kind you'd never let a machine send out on its own.
That gave me one rule I stuck to for the whole build: the agent gets no tool that commits anything. Everything else followed from that.
What it does
Rota is a shift scheduler for a small café. Fourteen staff, seven days, and an agent that works the rota with you rather than taking it off your hands.
When you open it, next week is about three quarters done. Fifteen slots are still empty, and there's one problem buried in it: Marco closes on Wednesday at half nine and opens again on Thursday at half six. Nine hours between shifts. The law wants eleven. Neither shift looks wrong on its own, which is why this kind of mistake survives a proofread.
Ask it to finish the rota and it makes about five tool calls. It reads the gaps, runs the solver, and stages fourteen assignments that add £1,384 to the wage bill. Then it stops and tells you the fifteenth slot can't be filled at all: Liam is seventeen so the curfew rules him out, and nobody else left is a certified shift lead.
Nothing on the calendar has moved. The fourteen proposals sit on top of the shifts they'd affect as dashed cards, and a drawer lists them one per line with a checkbox each. It shows you what approving would do to coverage, to your rule breaches, and to the cost. Take twelve of them and drop two if you like.
Then there's publishing, which is the one action that actually tells fourteen
people when they're working. I made that a declarative WebMCP form and left
toolautosubmit off it. Per the spec, that means an agent can fill the form in
and then has to stop. The browser puts focus on the submit button and waits for
a person. So the agent writes the note to staff, picks who gets notified, ticks
the confirmation box, and cannot press send.
A few other things that come out of building it this way:
- When it can't do something it tells you why and quotes the rule. "Everyone available is over their consecutive-day limit and Tom is the only other keyholder" is useful. "I couldn't find cover" isn't.
- Every tool call is logged with its arguments, what came back, how long it took, and who made it. You can export the lot as JSON and answer "why is Nadia on Friday night" properly.
- Undo works across approvals, because every edit carries a forward and a backward patch.
- No signup and no API key to try it.
How we built it
TypeScript, React and Vite. No backend, no database, nothing phoning home. The roster never leaves your tab, which is what makes doing this client side an honest choice instead of a shortcut.
Four reasons it's WebMCP and not an MCP server, roughly in the order they mattered to me:
The state is in the tab. The roster, the rules and the solver all run in the browser. A server would have to mirror the manager's working state, and half of that state is a proposal nobody has agreed to yet.
I didn't want the model reasoning about employment law. It never does. It calls
validate_schedule and repeats what the page's own engine said. When
assign_staff refuses, the refusal is a fact you can act on: "Liam Doyle is
not certified as Baker."
The agent needs to be able to point at things. When it says Thursday is the
problem it calls highlight and puts a ring round the two shifts on your
screen. Nothing running on a server can do that.
Some tools should only exist sometimes. Select a shift and two more appear,
scoped to it. Stage a proposal and revise_proposal shows up. Deselect and
they're withdrawn. Every one of those changes fires toolchange. A static tool
list on a server can't manage this, because a server has no idea what you're
looking at.
Under that there are three layers. The rules engine is plain TypeScript with no React and no DOM near it, because it's the part that has to be correct. Fifteen rules, split between the statutory and contractual ones and the softer ones about fairness, cost and preference.
Ranking candidates works by simulating and re-validating. To score someone for a shift I copy the roster, add the assignment, and run the real validator over the copy. It's not the fastest approach, but it means the ranking can never disagree with the rules, and that was worth more to me than the speed.
The solver is most-constrained-first, then two passes of local search. It's a heuristic and I'm not pretending otherwise. It won't always find the best answer, but it never returns an illegal rota and it's quick enough to run inside a tool call.
Both agent modes go through getTools() and executeTool(). The in-page agent
holds no references to its own tool functions on purpose, so my own UI is
exercising the same path an outside agent would take. There's also a polyfill
for document.modelContext, so the whole thing works in any browser today. If
you open it somewhere with real WebMCP the app notices and stands aside.
Challenges we ran into
The demo data was the hardest part, which I did not expect.
My first version of the broken week had 31 rule breaches and 65% coverage, and it just read as sloppy rather than realistic. Worse, I'd written more shifts than the staff had contracted hours to cover, so the solver couldn't finish no matter how good it was and filled nothing at all. Getting to a week that looks like a real half-finished rota, is genuinely fixable, and still leaves one slot that honestly can't be filled took several rounds of counting capacity against demand, role by role.
The edit model took two attempts. Applying patches straight meant two assignments to the same shift clobbered each other, so now each patch is rebased against the roster as it stands at the moment the edit lands.
And then there was the bug that nearly sank it. The "roster as proposed" value was being rebuilt every time it was read, and it's used as a state selector that compares by reference. So staging anything at all set off an infinite render loop, which tore down React and quietly unregistered all 32 tools. The app looked fine until the first time the agent tried to change something.
Accomplishments that we're proud of
That the constraint held. There are 36 tools and not one of them can commit, approve or publish. It would have been easy to add a commit tool for convenience halfway through and I didn't.
The correctness work is tested rather than claimed. 28 checks on the engine with no browser involved, 33 in a real browser, and 35 that click through the suggested prompts the way a judge would. They assert properties I actually care about: that anyone the ranker calls eligible really is eligible when you re-validate, that the solver never increases the number of hard breaches, that the cheap objective is never more expensive than the balanced one, that absence cover never backfills with the person who's absent, and that the publish form fills but refuses to submit.
And the polyfill, because it means anyone can open the URL and use every tool without an origin trial, an API key or a signup.
What we learned
Three things I got wrong first, all of which changed the code more than any prompt tuning did.
I had tools returning JSON. Switching them to write proper sentences was the
single biggest improvement in answer quality. The text half of a result is what
the model actually reasons over and then quotes at the user, so find_cover
now says "Mei Lin, score 88, no overtime, prefers evenings" instead of handing
back a tuple.
I was rejecting bad arguments. Models send "3" for an integer and "Thursday" for a date. Being strict about that is correct and useless, because the agent just burns a turn apologising. Now anything unambiguous gets coerced, and the coercion is written into the ledger so it isn't happening behind your back.
I was throwing exceptions. A model told "Marco is not certified as Baker" tries something else. One handed a rejected promise usually gives up. Failures come back as results with a sentence attached now.
The render loop bug taught me the other lesson, which is that a browser test earns its keep. Nothing in the type system or the unit tests would ever have caught it.
What's next for Rota: the agent proposes, you approve
Multiple weeks and multiple sites, since one week of one café is obviously a demo.
Shift swap requests from staff, as a second tool surface with less trust attached to it. Staff should be able to ask; only a manager should be able to agree.
And cross-origin federation, which is the part of the spec I most wanted to
reach and didn't. An agent running in an iframe would get at the venue's tools
through exposedTo, with a consent log showing the manager exactly what
crossed the boundary.
Built With
- css
- html
- javascript
- model-context-protocol
- openai
- react
- tailwindcss
- typescript
- vercel
- vite
- webmcp
Log in or sign up for Devpost to join the conversation.