Inspiration

Budgeting a household is a shared decision between people — and increasingly, between a person and an AI agent. Handing an agent full write access to your money is wrong. Making it scrape and click a human UI is worse. WebMCP offers a third path: give the agent real tools and keep the human in the loop at the only moment that matters — approving the change.

What it does

Budget Bench is a household budget workbench where the browser agent and the human work on the same live page:

  • The agent reads the whole budget with get_budget_summary (income, category budgets, spend, what's over/tight, pending proposals).
  • The agent acts through five proposing tools: find_savings (trims discretionary categories and moves the difference to savings), set_category_budget, add_category, record_spend, and propose_plan (full rebalance to hit a savings target, fixed costs untouched).
  • Every mutation is a proposal. Tool calls never write to the budget. Each one lands in the human-visible Proposal Tray with a plain-English summary and the agent's reason.
  • The human approves or rejects each proposal (or the whole tray). Approvals apply atomically; every proposal and every decision is written to the activity log.

Why WebMCP is the strong fit

  • Agent tools, not agent autonomy. Household money is exactly the case where an agent should propose and a human should decide. WebMCP tools carry structured schemas, so the agent gets a reliable contract instead of DOM scraping; the tray guarantees the human veto.
  • What people + agents can do together that was hard before: a two-minute conversation — "energy bill's up £35, find the money from discretionary" — becomes three reviewed, reversible, explained edits instead of a spreadsheet session. The same shared page also means the agent always sees the effect of its own proposals before the human decides.
  • Shared context, not replicated state. The page state is the app: no backend, no accounts, no auth replication. That's only possible because WebMCP runs tools where the state already lives — in the tab the human is looking at.

How we built it

One static page (vanilla HTML/CSS/JS, zero dependencies, zero build step) with a single state object and one render function. Six tools registered via document.modelContext.registerTool({name, description, inputSchema, execute}) per the WebMCP draft. Proposals carry apply closures, so approval is atomic and the tray doubles as a transaction log. Tool logic is covered by an automated Playwright suite (27 checks: registration, schema behaviour, proposal isolation, approve/reject application, error paths) which stubs only the browser side of modelContext — the app's tool code runs unchanged. Hosted on Surge. MIT licensed.

Challenges we ran into

Designing tools that are genuinely useful to an agent while guaranteeing the human veto: every mutating tool needed to become a proposal generator rather than a writer, which reshaped the whole state model. Also pacing propose_plan so fixed costs stay untouchable and the plan lands as ONE reviewable decision instead of a barrage of edits.

Accomplishments we're proud of

The proposal tray pattern: six tools, one hard invariant — the agent proposes, the human decides — and the activity log that proves it.

What we learned

That "agent-accessible" and "human-controlled" are not a trade-off if the tool layer is designed as a proposal system rather than a command API.

What's next for Budget Bench

CSV import of real bank data (local-only), per-category proposal bundles, and a "what-if" tool that compares two plans side by side in the tray.

Built With

Share this project:

Updates

Submission history