Inspiration

Six people go to Lisbon for ten days. One shared apartment, and everyone pays for random things along the way. At the end somebody has to work out who owes who, and it falls to one person, takes an evening, and gets argued about for a week.

Apps for this already exist and are widely installed. People still do not use them properly, and the reason is not the arithmetic. It is that logging one dinner means opening a form, typing an amount, choosing a payer, and ticking six checkboxes. Nobody does that on the last night of a holiday. So the ledger stays half empty and the group reconstructs it from memory.

The bottleneck was never the maths. It was the typing. That is a data entry problem, and it is exactly the shape of problem WebMCP is good at.

What it does

Even Split is a shared ledger for splitting costs across a group. Every expense knows who paid and who shares it, so a dinner somebody missed stays off their tab. Splits can be even, by shares, or exact amounts per person. A balance rail puts everyone on one scale, owing to the left and owed to the right. Any person opens to a line by line breakdown of how their balance was reached. A settlement panel works out the shortest set of payments that clears the whole group.

All of that works on its own, with no agent involved.

The page also registers 14 WebMCP site tools, so an agent can work the same ledger at the same time. Verified in the ChatGPT desktop app's built-in browser against the live deployment:

"Log these: I paid 180 for the apartment split six ways, dinner Tuesday was 90 that Sam paid but Priya wasn't there, and the taxi was 40 for me, Sam, Dana and Jo." It called list_people to orient itself, then made three separate add_expense calls with three different participant lists. That sentence cannot go into a form, because a form takes one expense with one list of people.

"Sam gave me 40 back in cash on Tuesday, and take Jo off the grocery run." record_payment and update_expense from one sentence. A four day old correction, and every downstream balance re-derived in view.

"Alex says he shouldn't pay full share for the groceries since he doesn't drink. Sort that out." No tool named, no amount, no split method. It found the right expense, inferred a weight change, and left Sam's existing double share alone.

"Why does Jo owe so much?" explain_balance returned the itemised trace and opened the same breakdown on the page. Nobody accepts a number from a chat message, so both sides read the same thing.

"Settle everything up." It staged five payments and stopped, reporting that nothing was recorded until the person approves.

How we built it

React and Vite, no backend. The ledger lives in localStorage and never leaves the browser.

The domain logic sits in one pure module: splits, balances, settlement, and explanation, with state in and values out, no React and no DOM. Every state change goes through a named action in a reducer.

src/lib/webmcp.js holds the whole tool layer and contains no business logic of its own. Each tool translates JSON Schema input into one of those named actions, which is the same action the interface dispatches when somebody clicks a button. One code path, one set of validation rules, two ways in.

person clicks "Add expense"  ─┐
                              ├─► actions.addExpense ─► reducer ─► render
agent calls add_expense      ─┘

Registration uses document.modelContext.registerTool with an AbortController, so tools unregister cleanly and never double register. Because WebMCP only runs in origin-isolated documents, the dev server and both deploy configs send Origin-Agent-Cluster: ?1.

Three details do the reliability work. Names resolve instead of ids, since an agent says "Priya" and not "p5", with "me" mapping to the current user and ambiguous input returning the real options rather than a guess. State is read through a ref, so a tool called ten minutes after registration sees the live ledger. And errors are written for the caller: a short exact split replies with how far off it is, an unknown name lists who is actually on the trip.

Money is integer cents throughout, with the largest remainder method for splits, so ten euros across three people comes out 3.34, 3.33, 3.33 and the shares always add back to the exact total.

Challenges we ran into

Making the agent's work visible. A change you cannot see is a change you cannot overrule, and that is the whole premise of co-presence. Five things now react on the page while an agent works: the beacon in the top bar names the running tool, an activity panel logs every call, the row that changed flashes, the balance panel is outlined when its figures are re-derived, and explain_balance opens the breakdown drawer so the agent can point at what it is discussing.

Stale closures. Tools register once, but they get called minutes later. Reading state directly at registration meant a tool would happily operate on a ledger that no longer existed. State is now read through a ref at call time.

Errors an agent can act on. The first version threw generic validation failures, and the agent would either give up or retry the same wrong thing. Rewriting every error message to name the actual problem and list the valid options let it correct itself without a round trip to the person.

Testing without a browser. Registration is the only part that needs a browser, so a headless script drives all 14 tools against the real reducer, including bad input, through a full realistic session. That turned a slow debugging loop into a fast one.

Deciding what not to give the agent. Every settlement is a money decision, so no tool can commit one.

Accomplishments that we're proud of

One sentence became three tool calls with three different participant lists. That was the central claim and it is demonstrated rather than asserted.

The agent inferred a split change from "he doesn't drink." No tool named, no amount, no method. It found the right expense, adjusted the weights, and noticed Sam already had a double share and left it alone.

Nothing in the tool surface can commit money. An agent stages a settlement and the payments appear in amber with nothing recorded. Approving is a button only the person can press, and the test suite fails if a tool matching commit or approve is ever added. When asked to settle up, the agent read that constraint out of the tool's return value and told the user to approve it on the page.

It refused to create duplicates. Asked to log three expenses that were already present, it checked existing state first and declined rather than blindly appending.

The maths is provably right. One suite asserts that every expense's shares sum to its total, that all balances sum to zero, and that applying a settlement brings every balance to zero.

What we learned

Tool descriptions are the product. More of the reliability came from rewriting descriptions than from writing code. The difference between an agent guessing and an agent orienting itself first was a sentence telling it to call list_people when it needs names.

Return the result, not an acknowledgement. Once every write replied with the per person shares and everyone's new balance, the agent started verifying its own work instead of assuming it had landed.

Errors are an interface. An error message is read by a model that can act on it, so it should name the problem and list the options.

WebMCP is not MCP with a different transport. The thing that matters is that the tools run in the page the person is looking at. MCP servers wrapping a hosted expense API already exist. None of them let somebody watch a balance re-derive and drag it back.

Being honest about prior art is stronger than claiming novelty. Newer split apps do parse a single spoken expense. Saying so, then drawing the real distinction, is more convincing than pretending the category is empty.

What's next for Even Split

Symmetry in the tool surface. There is add_person and no way to remove one, which an agent noticed by going looking for the button instead. The version worth building refuses with a reason: "Jo is on seven expenses and one payment, so I cannot remove them. Take Jo off those first?"

Receipts parsed in the browser. Local OCR would keep the no-server property while removing the last bit of manual entry.

Multiple currencies, which any real trip needs and which the integer cents model already accommodates.

The same architecture where the money is bigger. The propose-then-approve boundary and the itemised explanation matter more when reconciliation is an obligation rather than a social nicety: shared utilities across tenants, freelance collectives splitting project revenue, small business expense matching. The chore is the same shape and the stakes make the visible reasoning worth paying for.

Built With

Share this project:

Updates

Submission history