Inspiration
We counted the number of times in one week we typed some version of "context: I'm building a CRE tool, the stack is Python and SQLite, we decided on Clerk for auth last Tuesday" into a fresh chat window. It was eleven.
Not because the models are bad at remembering — because memory is a per-vendor feature. Claude remembers things in Claude. ChatGPT remembers things in ChatGPT. Cursor remembers nothing at all. The moment you move — and you move constantly, because different models are genuinely better at different jobs — your context stays behind. So you re-explain, or you keep a scratch doc and paste from it, which is just re-explaining with extra steps.
The friction isn't inside any one model. It's in the gaps between them. That's what we built for.
What it does
Vault is a memory layer that sits outside every assistant and serves all of them.
Ingest. Point it at Claude Code .jsonl sessions, a ChatGPT export, or plain Markdown
notes. It scrubs secrets, then runs extraction that pulls out durable facts — subject,
predicate, value, entity, category, confidence — plus a short episode summary for each
conversation.
Store. Facts land in SQLite with FTS5, and are mirrored as human-readable Markdown into an
XO Space project's memory/ folder. When a fact changes, the old one isn't overwritten — it's
marked superseded and moved under a History heading, so you can see that auth used to be
JWT before it became Clerk, and when that flipped.
Retrieve. vault ask "what did we decide about auth?" returns a packed context block under
700 tokens. A gate decides whether memory is even needed. A scope filter hard-limits results by
entity at the SQL level, so a prompt about project A cannot surface project B's facts, and
health or finance facts only appear when the prompt is actually about health or finance.
Serve. The same store is exposed as an MCP server (vault_catalog, vault_search,
vault_get, vault_remember), so any MCP-capable agent can query it mid-task, and as a
FastAPI web app for everything else.
Two rules we never bent: no transcript text ever enters memory/ — only a paraphrased evidence
string capped at 120 characters — and every injection is logged, so you can always answer
"what did it actually tell the model about me?"
How we built it
One codebase, two runtimes, split behind a single provider interface:
class Provider(Protocol):
name: str
def extract(self, conversation_text: str, schema: dict) -> dict: ...
def health(self) -> dict: ...
- Local mode uses the Anthropic API for extraction.
- Hosted mode uses an open-weight instruct model on a RunPod serverless endpoint via its OpenAI-compatible route, with the FastAPI app deployed on GalaxyGate. The endpoint scales to zero between requests, app state lives on a persistent path outside the container, and the browser never touches RunPod — only the server holds the key.
The build itself was run as an agent team in Claude Code. One orchestrator session wrote the fixed contracts first — Pydantic models, SQLite schema, the Markdown mirror, the provider interface — then ran three subagents in parallel against non-overlapping file ownership:
| Agent | Owns |
|---|---|
| librarian | ingest, scrub, extract, supersede |
| retriever | gate, scope, search, pack |
| qa | golden set, eval runner, leak test, lint config |
| integrator | CLI, MCP server, skill file, docs |
A second session on a hosted branch built the RunPod provider and deployed the web app, merged
back in the final phase. Ownership was written down before any agent started, which is the only
reason parallel edits didn't turn into merge conflicts.
Challenges we ran into
Scope is harder than search. Our first retrieval pass ranked well and leaked constantly — semantically similar facts from unrelated projects kept scoring high. Ranking can't be the thing that enforces isolation, because ranking is a suggestion. We moved scope to a hard SQL filter applied before ranking. Precision stopped being a probability.
Supersession vs. duplication. Same entity, same predicate, new value means the old fact is superseded. Same value means raise confidence and update the timestamp. Getting that wrong in either direction gives you a memory that either forgets decisions or accumulates five copies of the same one. Both paths are unit-tested.
Proving the no-leak rule. "We don't store transcripts" is worth nothing as a claim. We wrote
a test that asserts no 12-word span from anything in inbox/ appears anywhere in memory/. It
failed the first time — our episode summaries were quoting. That test is now the thing we'd
point a security-minded user at.
Token budgets are a packing problem. With a 700-token ceiling, you need an order (profile, then entity facts, then episodes), a rule for what degrades first (truncate an episode before dropping a fact), and deduplication against what this chat has already been told.
Model-agnostic JSON. A prompt that reliably returns schema-valid JSON from Claude does not automatically do so from an open-weight model. Both providers validate with Pydantic and retry once with the validation error appended — which is the agentic loop that matters most in practice.
[If hosted mode gave you deployment trouble — proxy, env vars, cold starts — add it here. Real specifics beat generic ones.]
Accomplishments that we're proud of
- Memory that is genuinely portable: ingested from one assistant, served to another, in one run.
- A scope guarantee we can demonstrate rather than assert, plus a leak test that enforces it.
- The same codebase running against a frontier API and an open-weight endpoint with a one-variable switch.
- It's auditable.
/auditshows every injection: what was sent, how many tokens, to which chat.
[Add your eval numbers here once vault eval has run — gate accuracy, recall@5, precision@5,
tokens mean and p95. Do not invent these.]
What we learned
- Contracts before parallelism. Freezing the schema and file ownership up front is what made four agents working simultaneously produce something that merged.
- Retrieval quality is a safety property, not just a UX one. Once a system has facts about your health and your finances in the same store as your project notes, "what gets surfaced" is the whole ballgame.
- An eval you run before every merge changes how you build. Numbers in
PROGRESS.mdmeant we could tell improvement from vibes.
What's next for Vault
- Hybrid retrieval: a RunPod embedding endpoint fused with BM25 by reciprocal rank fusion.
- A network volume for model weights, to cut cold starts.
- Encryption at rest and multi-device sync.
- Browser extension and a system-wide hotkey, so capture doesn't require a CLI.
- Shared team vaults, where scope rules become permissions.
Log in or sign up for Devpost to join the conversation.