Inspiration

At home I use excels to make my savings plan and at work I am the Agent Emperor. I dread coming into my financial planning tools and I wanted to supercharge it. When I saw WebMCP, I actually got interested.

I am a passionate programmer and read through the spec, I immediately wanted to build something to absorb the concepts. Thats why I build this tool. I would really want to be involved in defining the spec for WebMCP, not sure if I am allowed to or if I can participate? Am Beggingggggg.... begggingggg yoouuuu ohhh uuuuu...

Why is the usecase a strong fit for WebMCP

Getting an agent to help with a sensitive document normally means uploading the whole thing. With No Upload MCP users submit their confidential docs to the JS runtime in the browser, which deterministically parses PDF using text extraction - (for eg - bank statements) or in case of a pdf image uses OCR. This happens locally using user's CPU and browser, NO UPLOADs!!! Tadaa! Then the user can review extracted records, control what the agent may process, approve mappings, generate charts, and create editable financial or dietary plans without uploading their documents to an application server.
The agents via webmcp get access only to the required info and they supercharge it with generative context, on which human and agents can cross collaborate.

Earlier I could do all the above things, but I had to do searches separately , do a lot of copy paste and boring context switching. But now its quite immersive. I get everything done in one tab in one context.

Best part it runs in the browser for free. I get the best of both worlds - my privacy in my hands with a zero trust and all the capabilities of the agents.

Challenges I ran into

Convincing my wife to let me build this on a Saturday night

On a serious note, the documentation I felt was quite sparse. I wanted to try out scenarios where:

  1. The page could hook to the agent once an async flow was done, this seems to be not possible.
  2. Setting up certain guardrails for the agent and instructions. Maybe i didn't understand the spec well.

Accomplishments that I'm proud of

I don't need to use excel anymore. I can finally use AI without the concern of my private data leaving my machine!

What I learned

Loads about web MCP and I love the concept. I am already brimming with so many ideas for this! Thanks thanks thanks for running this hackathon. I already feel like a winner by building this use case for myself.

---------x-----------(below this line generated by AI reviewed by human)------------

WebMCP: What It Improves & Enables

✨ User Experience Enhancements

  • Visual Transparency: Delegate tasks in plain language and watch every consequence update instantly in the UI.
  • Explicit Merchant Rules: Converts raw descriptions (e.g., "amzn mktp" ➔ Amazon ➔ Groceries) into bulk-applicable rules rather than silent, row-by-row guesses.
  • Instant Rejections: Dismiss incorrect agent assumptions with a single click instead of an argument.
  • Strict "Propose, Don't Commit" Architecture: Mappings, ranges, and plans sit in a pending UI state with a visible source link. Nothing becomes durable just by looking plausible.

🤝 Advanced Human-Agent Collaboration

  • Real-Time Financial Planning: The agent drafts budget reductions (e.g., saving $50,000 for a car) from your actual spend. Drag a slider, type "keep my gym membership," and evaluate feasibility instantly.
  • Synchronous UI State: The agent reads the exact value under your cursor—not a turn from five minutes ago—because the UI mirrors inputs before writing to disk.
  • Interactive Health Reviews: Plot lab analytes against printed reference ranges, inject cited standards where data is missing, manually remove allergens from a drafted food plan, and have the agent re-verify the edited text before final publication.

🚫 Why Chat & Automation Fail

  • Chat interfaces force you to manually re-upload files or describe UI changes by hand.
  • Browser automation scrapes raw DOM objects and hopes they don't break.
  • Neither features a native contract for consent, provenance, or safe, staged execution.

WebMCP Architecture & Implementation

🛠️ Central Registry & Tool Dispatch

  • Centralized Management: Handled via src/core/mcp/registry.ts.
  • Schema Conversion: Converts typed internal tool specs into JSON input schemas.
  • Model Registration: Passes generated schemas directly to modelContext.registerTool.
  • Audited Execution: Routes every execute command through a secure dispatcher.
  • Detailed Logging: Records arguments, results, and duration for the in-page log.

🔄 Dynamic Tool Lifecycle

  • Baseline State: Eight core tools remain permanently registered.
  • Plugin Selection: Activating a plugin registers 17 finance tools or 12 labs tools.
  • Cleanup Mechanism: Retires inactive tools using their AbortController signal.
  • Schema Refresh: Automatically updates document-shaped schemas after data ingest.

📡 Resilient Host Discovery

  • Target Detection: Scans window, document, and navigator for modelContext.
  • Continuous Polling: Searches indefinitely since hosts can inject context long after load.
  • Session Longevity: Ensures hour-long persistent sessions never miss a late injection.

🛡️ Five-Tier Authority Model

Every tool declares a strict authority tier. This tier is explicitly built into the tool description provided to the model, acting as a public promise rather than simple internal bookkeeping:

  • Read: Standard data retrieval.
  • Attention: Modifies only the visible screen.
  • Consent: Restricts data return until a policy and grant check passes.
  • Staged: Proposes actions but commits no permanent changes.
  • Memory: Accesses page-owned durable state.

Built With

  • claude
  • codex
  • ocr
  • pdfjs
  • react
  • vite
  • webmcp
Share this project:

Updates

Submission history