Inspiration

It started with a familiar frustration. I save dozens of links every week — articles, posts, videos, threads — and revisit almost none of them. The bookmarks pile up, the tabs multiply, and most of what I carefully saved disappears into a graveyard of good intentions.

But the problem isn't saving. It's retrieval. And retrieval is exactly what AI is built for.

When I discovered WebMCP — an open standard that lets websites expose structured tools directly to AI agents — I realized I could bridge both sides of that gap. Not just build a smarter bookmark manager, but make my personal library natively accessible to any AI agent, without plugins, without API key sharing, without configuration. Just open the page, and your AI can search your library.

What I Built

SavedPocket is a personal AI library for saving and rediscovering content from across the web. It supports one-click saving via a Chrome extension, bulk import from WhatsApp chat history, and automatic AI analysis that generates a summary, tags, and category for every saved item.

The WebMCP layer exposes six tools at /webmcp:

  • search_library — semantic vector search across your entire saved library using natural language (finds results by meaning, not just keywords)
  • get_item — retrieve full details of a saved item including AI-generated summary, tags, extracted content, and notes
  • get_recent — list the most recently saved items in reverse chronological order
  • save_url — save a new URL directly to your library via AI agent; AI analysis and summarization run automatically after saving
  • list_collections — browse all curated collections in your library
  • get_collection_items — retrieve all items inside a specific collection

This means an AI agent can not only read your library but also write to it — asking ChatGPT "save this page for me" actually saves it to SavedPocket.

With SavedPocket open in a browser tab, ChatGPT can call these tools directly using WebMCP — no setup required on the AI side.

How I Built It

Stack: Next.js 15 (App Router), PostgreSQL with pgvector, Drizzle ORM, deployed on Render.

Semantic search runs entirely on-device using Xenova/multilingual-e5-small — a quantized ONNX model loaded via @huggingface/transformers. Embeddings are stored as 384-dimensional vectors in PostgreSQL and indexed with HNSW for sub-linear approximate nearest-neighbor search:

$$\text{similarity}(q, d) = 1 - \frac{q \cdot d}{|q| \cdot |d|}$$

Every saved item is embedded once and cached in the database. At query time, only a single embedding is computed (the user's message), then matched against stored vectors using the cosine distance operator <=>. This means search performance is independent of library size.

The WebMCP endpoint follows the MCP protocol over HTTP, handling initialize, tools/list, and tools/call methods. The Chrome extension additionally discovers WebMCP endpoints across the web using /.well-known/mcp.json, enriching saved items with structured content where available.

Challenges

Memory constraints on a shared server. The ONNX model takes ~50–60 MB of RAM,
and the first embedding inference on a cold server adds a transient spike. Running on a 512 MB instance caused OOM crashes early on. The fix was baking the model into the Docker image at build time and tuning Node.js GC with --max-old-space-size.

Idempotent startup seeding. The demo user and seed data are created automatically on every deploy via a startup hook. Getting this to be truly idempotent — handling existing users, collections, and items without constraint violations across deploys — required careful use of onConflictDoNothing and matching upserts to the actual database unique constraints.

WebMCP discovery in the Chrome extension. Fetching /.well-known/mcp.json from arbitrary origins introduces CORS and timing complexity inside a service worker. A 3-second timeout with graceful fallback to HTML scraping kept the extension reliable across sites that don't yet speak WebMCP.

What I Learned

WebMCP flips the model. Instead of AI agents scraping the open web or requiring
users to paste content into a chat window, websites can declare what they offer and let agents consume it cleanly. SavedPocket is a small demonstration of what that future looks like applied to personal knowledge — a library that any AI can search, the moment you're logged in.

The standard is early. Most sites don't expose WebMCP endpoints yet. But the infrastructure is there, and the pattern is simple enough that adoption can move quickly. I came away convinced this is how the agent web will actually work.

Built With

Share this project:

Updates

Submission history