Inspiration
Every AI assistant I've used has amnesia. Each conversation starts from zero — it doesn't remember my projects, my preferences, or what we discussed yesterday. Real personal assistance is impossible without memory. "Yaad" — the Hindi word for memory — is my answer: an always-on personal AI that truly remembers, learns my workflows, and gets smarter over time.
What it does
Yaad is a private personal AI with persistent memory and self-improving skills:
- Never forgets: bi-temporal memory — facts carry validFrom/validTo timestamps, so contradictions are resolved by retiring old facts, never deleting history. Hybrid RRF retrieval (dense embeddings + keyword + entity signals) scores recall@1 = 1.00 on our eval harness.
- Sees connections: automatic entity extraction builds a knowledge graph you can explore visually; background "dreaming" clusters memories into insights.
- Improves itself: when it spots a repeated workflow, it drafts a reusable SKILL.md pack — always approval-gated, never auto-applied.
- Researches deeply: Tavily-powered deep research runs planner → parallel searches → parallel page extraction → cited synthesis, plus fire-and-forget background research jobs with proactive "done" nudges.
- Acts proactively: reminders, morning briefings, and dream-insight nudges surface on their own.
- Shows its work: streaming responses, visible reasoning traces, parallel tool calls, sandboxed code execution, voice input, photo memory via vision model, PWA install, and Markdown chat export.
How we built it
- Brain: NVIDIA Nemotron models via Nebius Token Factory — Nemotron 3 Nano 30B for fast routing, Nemotron 3 Super 120B for deep reasoning, Qwen3-Embedding-8B for semantic memory, and a Nemotron vision model for photo memory.
- Stack: TypeScript throughout — Express backend, React + Tailwind frontend, JSON-based memory store with an MCP server exposing the same 7-operation interface to external agents.
- Deployed: frontend on GitHub Pages, backend on Render, with a real $0.50/day cost cap enforced server-side so the demo never burns budget.
Challenges we ran into
- The Render deploy failed in a sneaky way: a stale dashboard build command plus a stray backtick in the start command crashed every boot — fixed via the Render API.
- Our cost cap was silently dead (unmetered code paths returned null cost estimates). We revived it with conservative fallback pricing and per-turn spend reservations — now it genuinely trips before overspend.
- Free-tier reality: ephemeral disks (memory resets on redeploy) and cold starts — documented honestly and mitigated with a keepalive.
Accomplishments that we're proud of
- Memory eval: recall@1 = 1.00, MRR = 1.000 across 6 adversarial cases (contradictions, historical queries, distractors).
- A full 16-finding security audit fixed and verified — SSRF redirect-chain validation, rate limiting, sandboxed code execution.
- A real MCP server so other agents can use Yaad's memory.
What we learned
Memory isn't a feature you bolt on — it's the architecture. Bi-temporal validity, hybrid retrieval, and approval-gated self-improvement changed how we think about "personal" AI.
What's next for Yaad
Multi-user isolation with proper auth, persistent disk-backed memory, and a public skill marketplace where approved skill packs can be shared.
Built With
- embeddings
- express.js
- github-pages-nebius
- knowledge-graph
- model-context-protocol
- nebius
- node.js
- nvidia-nemotron
- pwa
- react
- render
- tailwind-css
- tavily
- typescript
- vite
Log in or sign up for Devpost to join the conversation.