Inspiration

Most "AI meal planning" demos are a chatbot that describes a plan in a text box - you still have to go copy the ingredients into a shopping list app yourself. WebMCP flips that: instead of an agent guessing at button selectors or a chat window narrating what you should do, a page can expose its actual capabilities as tools an agent calls directly. Meal planning turned out to be a good fit to prove that out. It's a genuinely tedious, recurring chore with real constraints - dietary restrictions, a grocery budget, food already sitting in the fridge about to go bad - the kind of multi-step, stateful task that's tedious for a human to do by hand and awkward for a chatbot to do by description, but a natural fit for an agent that can actually read and write app state.

What it does

PantryPilot is a meal planner and shopping list. A household sets dietary preferences (vegetarian, vegan, gluten-free, dairy-free, nut-free, or a free-form avoid-list), tracks what's in the pantry (including expiry dates), and optionally sets a weekly grocery budget. From there:

  • Search and rank a 16-recipe catalog by pantry match ("what can I make with what I already have"), by what's about to expire, by cost, by prep time, or by calories/protein.
  • Plan a single day, or hand the whole week to plan-week, which fills every empty day under whatever constraints you give it - tags, max prep time, a calorie/protein ceiling, the remaining budget, and a "use up what's expiring" priority - while never overriding an active dietary restriction and never repeating a recipe in the same week.
  • Get a shopping list generated automatically from the plan, aggregated across recipes and offset by what's already in the pantry.
  • See a per-day and weekly nutrition summary.
  • See exactly what happened and who did it - every action is logged and tagged human or agent - and undo the most recent change with one click, by either party.

All of that is exposed as 19 WebMCP tools registered via document.modelContext.registerTool(), so a human using the page and an agent acting on it are working in the same live session, not two separate experiences.

How we built it

Plain JavaScript, ES modules, no framework and no build step - the whole app is static files. The core design decision was a single shared domain module, js/actions.js, that owns every mutation (planning a recipe, updating the pantry, regenerating the shopping list, undo). Both the human-facing UI and every WebMCP tool's execute() call the exact same functions in that module, tagged with who's calling ("human" or "agent"), so there's no separate, more-powerful "agent path" that could drift out of sync with what a human can do. State lives in a small reactive store backed by localStorage; the UI and the tool layer both subscribe to its change event, so a change from either side re-renders the page immediately.

Since this session's sandbox had no WebMCP-capable browser to test against, we built an in-page "Agent Console" that drives the identical document.modelContext tool handlers through a plain form (pick a tool, edit the JSON arguments, run it) - the same code path a real agent would hit, just reachable without the Chrome 149 origin trial or ChatGPT's browser. It became a first-class part of the app rather than a throwaway test harness, since it also lets anyone verify every tool works without installing anything.

We also wrote a node:test suite (32 tests) against js/actions.js directly, covering dietary-conflict enforcement, shopping-list aggregation, budget math, and the undo system, and ran full Playwright end-to-end passes over every tab and every tool before each push.

Challenges we ran into

  • No native WebMCP runtime to test against. We couldn't rely on trying it in a real agent client during development, so correctness had to come from the domain layer itself (validation, tests, the Agent Console) - the tool registration is thin by design, and almost all of the actual logic lives in code we could run and test directly.
  • Keeping the agent from being more powerful than a human. It would have been easy to write a fast "agent path" that skips validation for convenience. Instead every tool routes through the same functions the UI uses, which meant designing actions.js first as a proper API, not as UI glue.
  • Undo without duplicating history. A single logical action like "add a recipe to Monday" cascades into a second internal update (the shopping list regenerating). A naive undo implementation would push a separate history entry for that cascade and undo the wrong thing. Snapshots are taken immediately before the actual mutating store.update() call inside each action, and the shopping-list regeneration that happens as a side effect of another action is marked silent and skipped by the snapshot - so one user-facing action is always exactly one undo step, verified with a dedicated test.
  • Shopping-list aggregation. Merging ingredient quantities across multiple recipes, scaling by servings, subtracting whatever's already in the pantry, and preserving which items are already checked off across a regeneration all had to work together without double-counting or losing state.

Accomplishments that we're proud of

  • 19 WebMCP tools that are genuinely load-bearing - plan-week does real constraint-based selection (dietary safety, no repeats, budget, expiry priority), not a lookup table dressed up as a tool.
  • An architecture where an agent provably can't do anything a human couldn't do by clicking around, because they're the same code.
  • A working undo system - a real implementation of the reversibility idea in WebMCP's own security guidance, not just a line in the description.
  • 32 passing unit tests plus a full Playwright pass over every tab and every tool, run before every push.
  • Zero backend, zero build step, deployable to Netlify, Vercel, or Cloudflare Pages in minutes.

What we learned

  • WebMCP's imperative API is small on purpose - registerTool({ name, description, inputSchema, execute }) is most of it - which pushes the real design work down into the domain logic underneath, where it belongs.
  • Tools that compose (plan-week reusing search-recipes's own ranking logic) produce much better agent behavior than a flat pile of unrelated CRUD tools would.
  • Trust in a page an agent can write to has to be built into the architecture - one shared code path, validation at the domain layer, and reversibility - not asserted after the fact in a description.

What's next for PantryPilot

  • Testing against a real WebMCP origin trial browser and ChatGPT's browser once broadly available, beyond the in-page Agent Console.
  • Redo (undo is currently one level deep, with no way back).
  • Recipe ratings/favorites and multi-week history ("plan like last week").
  • A substitution tool that suggests a swap when an ingredient is missing or newly avoided.
  • Optional lightweight sync so a plan isn't tied to one browser's localStorage.
  • Exploring WebMCP's declarative API and cross-origin exposedTo, e.g. a grocery-delivery origin reading the shopping list directly.

Built With

  • agent
  • mcp
  • model-context-protocol
  • netlify
  • webmcp
Share this project:

Updates

Submission history