The problem
A year of my bank statement is 154 pages and 1,630 rows. I have never read it. Nobody reads it. So the standing instructions carry on, the money goes roughly where I half-remember it going, and anything worth a second look sits unread on page 91.
Uploading that PDF to a chatbot does not fix it. The model reads what it can, skips what it cannot, and quietly miscounts the rest. And every figure in a bank statement is load-bearing.
What Passbook does
Passbook parses the statement inside the browser tab, reconciles every row against the statement's own printed running balance, and then publishes the result as WebMCP tools. Your own agent asks the questions; the page does the arithmetic.
Built on real statements, not fixtures:
| Real HDFC statement | 154 pages, 1,630 rows |
| Parsed | 1,630 of 1,630, zero failures, 100% bank-reference coverage |
| Running-balance chain | intact end to end |
| Also parsed | Kotak Mahindra (107 rows), RBL (143), plus a CSV path for any bank |
| Fixed-amount recurring merchants found | 10 |
| Candidate pairs worth a second look | 9, out of 1,630 rows |
| Of those, confirmed bank errors | 0. Every one was a payment I meant to make |
The part I want to be honest about
An earlier version of this project led with "it found the money you already lost". It detected pairs of same-amount charges to the same counterparty within a few days, and called them duplicates.
Then I audited all nine against my own memory of the account. Every single one was intentional. The postings are real: the balance chain reconciles across every pair, and four are adjacent rows where the balance falls by that amount twice in succession. So the detector was not inventing anything. It simply cannot see intent, and no ledger can.
So Passbook does not claim to find duplicate charges. It claims to read 1,630 rows and hand back the nine worth checking, and to say plainly that only the account holder can settle them. A tool that told me it had found my money would have been wrong nine times out of nine.
Precision as an error detector: 0 of 9. Usefulness as a filter: 1,630 rows down to nine I can check in five minutes.
What it actually gives you
Where the money went. Counterparties ranked by spend, how many payments each, and what share of everything that left the account they represent. Pure arithmetic over reconciled rows.
What leaves before you look. Standing commitments that debit an identical amount repeatedly, annualised from the cadence actually observed rather than an assumed monthly one. The row shows what leaves each time; the number that changes your mind is what it comes to in a year, and nobody does that multiplication while reading a PDF.
The few worth checking, with their evidence: both dates, both bank references, and
why a reversal was ruled out. A naive (date, amount, merchant) match returns 367 pairs
on this statement: 222 from one investment platform, 110 from another, all instalment
plans. Nine survive. Getting from 367 to 9 is the work.
And you can just ask. "How much did I spend on food?" The agent decides which
merchants count as food; the total_spent tool does the summing over reconciled rows and
returns the terms it counted. The number is computed, not a model's mental arithmetic.
Why WebMCP is a strong fit here
A bank statement is the worst possible thing to hand a model: too long to read, and every figure matters. WebMCP inverts it. The page keeps the ledger, parses it deterministically, reconciles each row against the printed balance, and exposes questions an agent can ask rather than a document it has to skim.
The agent brings intent and language. The page brings arithmetic that is either right or fails a checksum. Neither does this alone.
What this is a prototype of
I built this outside a bank, on a PDF I downloaded. But the interesting version is inside one.
Banks are never going to let an agent move money. That conversation is over before it starts, and it should be. Reading is a different question entirely, and reading is where almost all of the value is. Nobody needs an agent to make a payment; they need one to tell them where sixteen hundred payments went.
Passbook is that surface, built to the shape a bank could actually ship:
- Seven of its nine tools carry
readOnlyHint. The two that write, write to a document in the reader's own browser. There is no network write anywhere in the tool layer, and no endpoint that could move a rupee. The tool list declares this before an agent calls anything. - The trust boundary already exists. A customer logged into their bank is already authenticated, already inside a session the bank controls, already looking at their own data. WebMCP makes that existing page agent-readable without minting a token for a third party, without publishing a new public API, and without a bulk export leaving the building. The agent never receives a credential. It receives the fields a tool chose to return, inside a session the customer opened.
- The surface can be a function of state, so a tool that does not apply is not registered rather than registered-and-refusing.
- Everything the agent did is on the page. The Activity log lists the exact field set every call returned, which is the kind of artefact a regulated institution has to be able to produce anyway.
Swap the parsed PDF for the live ledger and the same nine tools work unchanged. That is the point of doing it on a real 1,630-row statement rather than a fixture: the tool shape is proven against the messiest version of the data before anyone has to trust it with the clean one.
How it improves the experience
You stop reading. A year of transactions becomes something you interrogate in your own words and get an answer computed over reconciled rows. And the agent receives only the field set a tool chose to return. The Activity panel lists those exact fields, per call, so the minimisation is something you can check rather than something I claim.
What the human and the agent can now do together
The page can narrow 1,630 rows to nine worth a look, and prove they are worth it. What it cannot do is know whether you meant to pay twice.
So get_duplicate_candidates returns those findings with a question attached.
Passbook has the ledger; it does not have the account holder's memory. The agent puts it
to you in its own words, your answer settles the case, and it is recorded with your
reason. The same question appears on screen, in the same words, for anyone using the page
by hand.
A ledger that asks, a person who answers, and a document neither wrote alone.
Every tool has a human equivalent. Anything the agent can do, you can do by clicking. A browser without WebMCP loses the agent, not the product.
The WebMCP implementation
Nine tools, and the registered set is a function of application state rather than a fixed
catalogue. Before a statement is imported the analysis tools are not registered at all.
Drafting withdraws once every candidate is handled. A tool that exists and returns "you
cannot use me yet" asks the model not to do something it can still do, so Passbook does
not ask, and the page ships a live ablation of that difference at ?ablation=instruction.
Read tools carry readOnlyHint. Anything returning statement narration carries
untrustedContentHint, because that text was written by whoever sent the money and is not
Passbook's. Neither hint changes behaviour; both are declarations to the agent.
Several details are measurements rather than assumptions:
- Registration needs only
registerTool. An earlier version demandedgetToolsandexecuteTooltoo, which would have refused every registration in a browser that exposes only the one that matters, silently, in the exact environment this is built for. executeTool's argument type is not portable. Chrome wants a JSON string and rejects an object; an agent's in-app browser wants an object and rejects a string. The registry negotiates the form and caches it, and only retries when the error is an input-shape complaint, because retrying a mutating tool on any other error could draft twice.- Gemini 3 rejects a follow-up that omits the
thoughtSignatureit attached to afunctionCall, so the built-in agent could call one tool and then die. Each turn now keeps the provider's own representation of its reply and replays it unchanged. - Chrome fired zero
toolchangeevents for a page changing its own tool map, so the UI reads the surface back throughgetTools()and never trusts the event. - OpenAI's API sends no CORS headers, so no web page can call it. Measured, not assumed: Gemini answers 400 to an invalid key, Anthropic 401, OpenAI fails with
TypeError: Failed to fetch.
On privacy, precisely
Parsing happens in the tab. The agent receives only the fields a tool returned, and the Activity panel lists them per call. That is data minimisation, not secrecy. Tool results do reach the model, and I would rather say so. Statement passwords go straight to the pdf.js worker and are never stored or logged. No real statement is in the repository.
Honest limits
- Passbook does not move money. It produces a document you send to your bank.
- The page is not a security boundary against its own user. Enforcement here is about what the model can do, not about defeating DevTools.
- Chrome's own guidance says it is impossible to guarantee safety inside a large language model. I make no claim to have solved prompt injection.
- Detection is a heuristic with stated confidence. It surfaces candidates for a human to judge; it does not decide.
Built With
- framer-motion
- netlify
- pdf.js
- react
- tailwind-css
- typescript
- vite
- vitest
- webmcp
Log in or sign up for Devpost to join the conversation.