Inspiration

I didn't want to build another "AI does your admin" demo. They're easy to make and they all feel the same, because the interesting part gets skipped. An agent doing a job for you is only half a product. The other half is how you check it.

Shift rotas turned out to be a good place to ask that question. A rota isn't really a spreadsheet, it's a promise to fourteen people about their week. Get it wrong and someone's childcare falls through. So it's exactly the kind of work you'd happily hand to a machine, and exactly the kind you'd never let a machine send out on its own.

That gave me one rule I stuck to for the whole build: the agent gets no tool that commits anything. Everything else followed from that.

What it does

Rota is a shift scheduler for a small café. Fourteen staff, seven days, and an agent that works the rota with you rather than taking it off your hands.

When you open it, next week is about three quarters done. Fifteen slots are still empty, and there's one problem buried in it: Marco closes on Wednesday at half nine and opens again on Thursday at half six. Nine hours between shifts. The law wants eleven. Neither shift looks wrong on its own, which is why this kind of mistake survives a proofread.

Ask it to finish the rota and it makes about five tool calls. It reads the gaps, runs the solver, and stages fourteen assignments that add £1,384 to the wage bill. Then it stops and tells you the fifteenth slot can't be filled at all: Liam is seventeen so the curfew rules him out, and nobody else left is a certified shift lead.

Nothing on the calendar has moved. The fourteen proposals sit on top of the shifts they'd affect as dashed cards, and a drawer lists them one per line with a checkbox each. It shows you what approving would do to coverage, to your rule breaches, and to the cost. Take twelve of them and drop two if you like.

Then there's publishing, which is the one action that actually tells fourteen people when they're working. I made that a declarative WebMCP form and left toolautosubmit off it. Per the spec, that means an agent can fill the form in and then has to stop. The browser puts focus on the submit button and waits for a person. So the agent writes the note to staff, picks who gets notified, ticks the confirmation box, and cannot press send.

A few other things that come out of building it this way:

  • When it can't do something it tells you why and quotes the rule. "Everyone available is over their consecutive-day limit and Tom is the only other keyholder" is useful. "I couldn't find cover" isn't.
  • Every tool call is logged with its arguments, what came back, how long it took, and who made it. You can export the lot as JSON and answer "why is Nadia on Friday night" properly.
  • Undo works across approvals, because every edit carries a forward and a backward patch.
  • No signup and no API key to try it.

How we built it

TypeScript, React and Vite. No backend, no database, nothing phoning home. The roster never leaves your tab, which is what makes doing this client side an honest choice instead of a shortcut.

Four reasons it's WebMCP and not an MCP server, roughly in the order they mattered to me:

The state is in the tab. The roster, the rules and the solver all run in the browser. A server would have to mirror the manager's working state, and half of that state is a proposal nobody has agreed to yet.

I didn't want the model reasoning about employment law. It never does. It calls validate_schedule and repeats what the page's own engine said. When assign_staff refuses, the refusal is a fact you can act on: "Liam Doyle is not certified as Baker."

The agent needs to be able to point at things. When it says Thursday is the problem it calls highlight and puts a ring round the two shifts on your screen. Nothing running on a server can do that.

Some tools should only exist sometimes. Select a shift and two more appear, scoped to it. Stage a proposal and revise_proposal shows up. Deselect and they're withdrawn. Every one of those changes fires toolchange. A static tool list on a server can't manage this, because a server has no idea what you're looking at.

Under that there are three layers. The rules engine is plain TypeScript with no React and no DOM near it, because it's the part that has to be correct. Fifteen rules, split between the statutory and contractual ones and the softer ones about fairness, cost and preference.

Ranking candidates works by simulating and re-validating. To score someone for a shift I copy the roster, add the assignment, and run the real validator over the copy. It's not the fastest approach, but it means the ranking can never disagree with the rules, and that was worth more to me than the speed.

The solver is most-constrained-first, then two passes of local search. It's a heuristic and I'm not pretending otherwise. It won't always find the best answer, but it never returns an illegal rota and it's quick enough to run inside a tool call.

Both agent modes go through getTools() and executeTool(). The in-page agent holds no references to its own tool functions on purpose, so my own UI is exercising the same path an outside agent would take. There's also a polyfill for document.modelContext, so the whole thing works in any browser today. If you open it somewhere with real WebMCP the app notices and stands aside.

Challenges we ran into

The demo data was the hardest part, which I did not expect.

My first version of the broken week had 31 rule breaches and 65% coverage, and it just read as sloppy rather than realistic. Worse, I'd written more shifts than the staff had contracted hours to cover, so the solver couldn't finish no matter how good it was and filled nothing at all. Getting to a week that looks like a real half-finished rota, is genuinely fixable, and still leaves one slot that honestly can't be filled took several rounds of counting capacity against demand, role by role.

The edit model took two attempts. Applying patches straight meant two assignments to the same shift clobbered each other, so now each patch is rebased against the roster as it stands at the moment the edit lands.

And then there was the bug that nearly sank it. The "roster as proposed" value was being rebuilt every time it was read, and it's used as a state selector that compares by reference. So staging anything at all set off an infinite render loop, which tore down React and quietly unregistered all 32 tools. The app looked fine until the first time the agent tried to change something.

Accomplishments that we're proud of

That the constraint held. There are 36 tools and not one of them can commit, approve or publish. It would have been easy to add a commit tool for convenience halfway through and I didn't.

The correctness work is tested rather than claimed. 28 checks on the engine with no browser involved, 33 in a real browser, and 35 that click through the suggested prompts the way a judge would. They assert properties I actually care about: that anyone the ranker calls eligible really is eligible when you re-validate, that the solver never increases the number of hard breaches, that the cheap objective is never more expensive than the balanced one, that absence cover never backfills with the person who's absent, and that the publish form fills but refuses to submit.

And the polyfill, because it means anyone can open the URL and use every tool without an origin trial, an API key or a signup.

What we learned

Three things I got wrong first, all of which changed the code more than any prompt tuning did.

I had tools returning JSON. Switching them to write proper sentences was the single biggest improvement in answer quality. The text half of a result is what the model actually reasons over and then quotes at the user, so find_cover now says "Mei Lin, score 88, no overtime, prefers evenings" instead of handing back a tuple.

I was rejecting bad arguments. Models send "3" for an integer and "Thursday" for a date. Being strict about that is correct and useless, because the agent just burns a turn apologising. Now anything unambiguous gets coerced, and the coercion is written into the ledger so it isn't happening behind your back.

I was throwing exceptions. A model told "Marco is not certified as Baker" tries something else. One handed a rejected promise usually gives up. Failures come back as results with a sentence attached now.

The render loop bug taught me the other lesson, which is that a browser test earns its keep. Nothing in the type system or the unit tests would ever have caught it.

What's next for Rota: the agent proposes, you approve

Multiple weeks and multiple sites, since one week of one café is obviously a demo.

Shift swap requests from staff, as a second tool surface with less trust attached to it. Staff should be able to ask; only a manager should be able to agree.

And cross-origin federation, which is the part of the spec I most wanted to reach and didn't. An agent running in an iframe would get at the venue's tools through exposedTo, with a consent log showing the manager exactly what crossed the boundary.

Built With

Share this project:

Updates

Submission history