Paste this newer version:

# nmemo  
### _Context that helps AI agents start informed._

> **What if an AI agent could understand your project before you had to explain it again?**

## The problem

Every project creates context everywhere:

- GitHub has the code
- PDFs hold research and decisions
- Notes contain the “why”
- Previous AI sessions contain unfinished work

But when someone asks an AI agent a question, the agent usually sees only a small fragment of that story.

The result is familiar: repeated explanations, irrelevant answers, wasted LLM tokens, and agents that confidently answer without enough evidence.

For student teams, community projects, and solo builders, this is especially painful. Small teams move fast, knowledge changes often, and important decisions easily disappear between people, tools, and coding sessions.

## The idea

**nmemo is a context orchestration layer for AI agents.**

It connects workspace knowledge and decides what an agent should know *before* the agent responds.

Instead of dumping every document into a model prompt, nmemo creates a compact context package:

```text
Question
  -> retrieve relevant sources
  -> rank evidence
  -> remove duplicates
  -> handle conflicts
  -> apply token budget
  -> return citations
  -> LLM answer

The goal is simple:

Give agents the right context, not all context.

How it works

A user connects sources such as:

  • Documents and PDFs
  • GitHub repositories
  • Workspace memory
  • Future team integrations such as Notion and Slack

When an agent asks a question, nmemo retrieves evidence from those sources and measures how each one performs.

The Playground makes this visible. It shows:

  • Source latency — how long Memory, GitHub, and document retrieval took
  • Ranking scores — why stronger evidence is selected first
  • Grounding — how much of the response is supported by retrieved context
  • Token usage — exactly how much context each source consumes
  • Citations — where the final answer came from

This turns prompt context from a black box into an explainable system.

Why this is a distributed systems project

nmemo coordinates information from independent services that have different latency, availability, and data formats.

A single request can travel through:

Vercel dashboard
  -> Render API
  -> PostgreSQL workspace state
  -> Qdrant document retrieval
  -> Voyage embeddings
  -> nmemo orchestration engine
  -> Groq LLM response

The hard part is not calling an LLM.

The hard part is deciding what happens when one source is slow, when two sources repeat the same information, when information conflicts, or when the LLM only has room for a limited number of tokens.

nmemo treats context as a bounded and observable resource.

How I built it

I built nmemo as a full-stack TypeScript application.

  • Next.js powers the dashboard
  • Clerk protects user sessions
  • Express provides the API layer
  • PostgreSQL stores workspace and connector data
  • Qdrant handles semantic document retrieval
  • Voyage AI creates embeddings
  • Groq generates LLM responses
  • Vercel hosts the frontend
  • Render hosts the API

I also built an SDK so another AI application can call nmemo before it calls its own model:

const context = await nmemo.getContext({
  query: "What decisions were made for this project?",
  tokenBudget: 4000,
});

Challenges

The biggest challenge was handling multiple sources without treating them as equal.

A PDF, GitHub repository, and memory entry can all return information, but they may be slow, duplicated, outdated, or only partly relevant.

I solved this by making the orchestration decision visible:

  1. Retrieve source evidence
  2. Measure source latency
  3. Rank by relevance and reliability
  4. Remove duplicate context
  5. Fit the best evidence inside a token budget
  6. Return citations with the final context

What I learned

I learned that vector search alone is not enough for reliable AI agents.

Agents need a system that can decide:

  • What is relevant?
  • What is trustworthy?
  • What is repeated?
  • What fits inside the context window?
  • Where did the answer come from?

What’s next

I want to expand nmemo with more workspace connectors, graceful fallback when individual sources are unavailable, portable context packages for coding agents, and better shared memory for teams.

nmemo is not another chatbot.

It is the context decision layer that helps AI agents know the right thing before they answer.

Built With

Share this project:

Updates

Submission history