Inspiration
Most autonomous agents and LLM frameworks suffer from amnesia: the moment a session ends or the context window overflows, critical user preferences and episodic memory vanish. Existing memory architectures force developers into proprietary cloud vector databases—introducing latency, privacy risks, and ongoing SaaS costs. We set out to build a production-grade, zero-cloud alternative that runs entirely on standard consumer hardware.
What it does
RecallOS gives AI agents persistent, continuous long-term memory running 100% locally with sub-12ms latency:
Hybrid Retrieval: Merges dense vector embeddings (FAISS/NumPy) with sparse lexical search (SQLite FTS5 / BM25), recency decay, and importance weighting.
Causal Knowledge Graph: Extracts entity relationships (RELATED_TO) and automatically resolves memory conflicts (SUPERSEDES).
Autonomous Fact Extraction: Converts raw dialogue into structured memory blocks using local LLMs (Qwen 2.5 via Ollama) with instant rule-based fallback.
Token-Budgeted Context Builder: Deduplicates, ranks, and injects optimal context into prompt windows without exceeding strict token limits.
Interactive 3D Dashboard: Features a dark-mode Next.js 14 console with a 360° interactive spatial memory engine visualizer.
Universal Interfaces: Ships with a Python SDK (recallos-sdk), FastAPI REST backend, and a native Model Context Protocol (MCP) server for Claude Desktop.
How we built it
Core & Indexing: Python, SQLite (FTS5 full-text indexing), FAISS, and NumPy for vector similarity calculations.
Inference Runtime: Local Ollama runtime running Qwen 2.5 models for autonomous memory extraction and graph derivation.
API & Integrations: FastAPI for high-throughput REST endpoints and native Anthropic MCP server implementation.
Frontend: Next.js 14, React, Tailwind CSS, and custom 3D spatial visualization components.
Challenges we ran into
Local Latency Optimization: Balancing dense vector retrieval, lexical BM25 filtering, and graph traversals under a 15ms threshold on CPU-only machines. We solved this with optimized in-memory NumPy/FAISS operations and tuned SQLite FTS5 queries.
Memory Drift & Conflict Resolution: Preventing contradictory statements from hallucinating agent contexts. We engineered an automated supersession mechanism that dynamically detects outdated facts and updates edge dependencies.
Accomplishments that we're proud of
Achieved 12.26ms average end-to-end retrieval latency and 100% Recall@5 on an 8-core consumer laptop.
Zero remote network calls: zero telemetry, zero data leakage, and zero cloud subscription overhead.
Native MCP support allowing Claude Desktop and Cursor to leverage persistent local memory out of the box.
What we learned
Local-first AI architectures can match or exceed cloud-hosted memory performance when dense semantics and sparse lexical techniques are tightly coupled with causal graphs.
What's next for RecallOS: Local Memory Operating System for AI Agents
Multi-agent collaborative memory pooling.
Native TypeScript/Node.js SDK.
Distributed synchronization option using PostgreSQL/pgvector plugins for hybrid edge-cloud deployments.


Log in or sign up for Devpost to join the conversation.