Inspiration
Budgeting a household is a shared decision between people — and increasingly, between a person and an AI agent. Handing an agent full write access to your money is wrong. Making it scrape and click a human UI is worse. WebMCP offers a third path: give the agent real tools and keep the human in the loop at the only moment that matters — approving the change.
What it does
Budget Bench is a household budget workbench where the browser agent and the human work on the same live page:
- The agent reads the whole budget with
get_budget_summary(income, category budgets, spend, what's over/tight, pending proposals). - The agent acts through five proposing tools:
find_savings(trims discretionary categories and moves the difference to savings),set_category_budget,add_category,record_spend, andpropose_plan(full rebalance to hit a savings target, fixed costs untouched). - Every mutation is a proposal. Tool calls never write to the budget. Each one lands in the human-visible Proposal Tray with a plain-English summary and the agent's reason.
- The human approves or rejects each proposal (or the whole tray). Approvals apply atomically; every proposal and every decision is written to the activity log.
Why WebMCP is the strong fit
- Agent tools, not agent autonomy. Household money is exactly the case where an agent should propose and a human should decide. WebMCP tools carry structured schemas, so the agent gets a reliable contract instead of DOM scraping; the tray guarantees the human veto.
- What people + agents can do together that was hard before: a two-minute conversation — "energy bill's up £35, find the money from discretionary" — becomes three reviewed, reversible, explained edits instead of a spreadsheet session. The same shared page also means the agent always sees the effect of its own proposals before the human decides.
- Shared context, not replicated state. The page state is the app: no backend, no accounts, no auth replication. That's only possible because WebMCP runs tools where the state already lives — in the tab the human is looking at.
How we built it
One static page (vanilla HTML/CSS/JS, zero dependencies, zero build step) with a single
state object and one render function. Six tools registered via
document.modelContext.registerTool({name, description, inputSchema, execute}) per the
WebMCP draft. Proposals carry apply closures, so approval is atomic and the tray doubles
as a transaction log. Tool logic is covered by an automated Playwright suite (27 checks:
registration, schema behaviour, proposal isolation, approve/reject application, error
paths) which stubs only the browser side of modelContext — the app's tool code runs
unchanged. Hosted on Surge. MIT licensed.
Challenges we ran into
Designing tools that are genuinely useful to an agent while guaranteeing the human veto:
every mutating tool needed to become a proposal generator rather than a writer, which
reshaped the whole state model. Also pacing propose_plan so fixed costs stay untouchable
and the plan lands as ONE reviewable decision instead of a barrage of edits.
Accomplishments we're proud of
The proposal tray pattern: six tools, one hard invariant — the agent proposes, the human decides — and the activity log that proves it.
What we learned
That "agent-accessible" and "human-controlled" are not a trade-off if the tool layer is designed as a proposal system rather than a command API.
What's next for Budget Bench
CSV import of real bank data (local-only), per-category proposal bundles, and a "what-if" tool that compares two plans side by side in the tray.
Built With
- ai-agents
- budget
- html
- javascript
- webmcp
Log in or sign up for Devpost to join the conversation.