Inspiration
Most "agent memory" demos fix forgetting. We went after the opposite failure: an agent that remembers everything, weights it all equally, and confidently acts on preferences you abandoned months ago. An agent that recalls nothing gives you a generic answer and you correct it. An agent with flat memory gives you a confidently personalized answer that's a year out of date — and you don't catch it, because it sounds like it knows you.
Here travel-agent project with a corpus of real landmark data, so we asked: what happens when the traveler changes and the memory doesn't?
What it does
A travel planner backed by eleven months of one traveler's decisions in Elasticsearch. Four of those decisions reversed:
| Stale original | Age | Replacement | Age |
|---|---|---|---|
| Walk the entire city, 20k steps/day | 330d | Knee injury — 3-mile cap, transit-adjacent | 45d |
| Museum-first itineraries (SFMOMA) | 300d | Stop routing me through museums | 25d |
| Pre-book Alcatraz and the Zoo | 280d | Skip ticketed attractions | 35d |
| Late shows at The Fillmore | 250d | Nothing scheduled past 9pm | 20d |
Ask it to plan a free day in San Francisco. With flat memory it books an 8pm show and a 20,000-step walking route — for someone with a torn knee. With recency-weighted recall it returns a transit-adjacent, ticket-free, museum-free day that ends in bed by 9.
Nothing is deleted and nothing is edited. Supersession is a retrieval problem, not a storage problem.
How we built it
Mastra for the agent loop, with remember and recall as tools the model chooses to
call — visible as explicit steps in the Studio trace rather than hidden context stuffing.
Elasticsearch as the memory store, with semantic_text computing embeddings
server-side at index time. Recall is one ES|QL query:
FORKruns two branches in parallel — BM25 over title/content/tags, and semantic search over the embedded fieldsFUSE LINEARmerges them at 0.3 BM25 / 0.7 semantic, because travel preference is prose rather than identifiersDECAY(created_at, NOW(), 1080 hours)multiplies the fused relevance score by recency
Data: 20 real San Francisco landmarks from an existing travel-agent project, with a 15-entry episodic memory layer authored on top spanning 3 to 330 days.
Challenges we ran into
The dataset had to be built to make the wrong answer win on merit. Our first instinct was to write short corrections. That produced a demo where decay changed the ranking from slightly-right to more-right — undramatic and unconvincing. So we wrote the superseded entries ~3x longer than their replacements (749 chars vs 224), packed with specific rationale. Now BM25 and the embeddings both legitimately prefer the stale memory, and only time-weighting can overturn it.
Retrieval was correct while the answer was still wrong. The agent recommended a Twin
Peaks hike to a user with a knee injury. The mobility memory ranked 6th; the tool's default
limit was 5. Retrieval quality and retrieval quantity turned out to be separate
failure modes — we raised the limit to 8.
The LLM kept rescuing bad retrieval, which broke our demo. With decay off, the agent often reasoned its way out of the contradiction on its own. That made the before/after non-deterministic and unusable as a live demo. We moved the demo down to the retrieval layer, which is pure ES|QL and reproducible to four decimal places.
Accomplishments that we're proud of
We measured our own failure rate instead of asserting a win. With decay off, the agent resolved three of the four reversals correctly by reasoning over the retrieved contradictions. It missed the fourth and booked a concert the night before a hackathon.
That number is the actual argument for this project: reasoning over stale memory works until it doesn't, and you can't tell which run you're getting. Decay at retrieval time makes correctness structural instead of lucky.
We can also show the window is a true half-life from our own data — the 45-day memory scores 0.7750 undecayed and 0.3899 at a 45-day window. Exactly half.
What we learned
- Decay multiplies relevance rather than replacing it. Our 5-day trip context stayed ranked #1 in both conditions because it was recent and relevant. Decay is not recency bias.
DECAY's third argument must be atime_duration(1080 hours), not adate_period.semantic_texton Serverless removes the entire embedding pipeline — no model choice, no vector generation, no client-side embedder.- Tuning is not just the knobs. The decay window, the fusion weights, the result limit, and the agent's conflict-resolution instructions are four separate levers, and the limit was the one that actually broke us.
- Storing the age as a real timestamp rather than a staleness flag is what makes the result honest — there is nothing in the index marking a memory as superseded for the query to key off.
What's next for Superseded - episodic memory for travel planning
- Write-path in the demo. The
remembertool already works; capturing a reversal live and watching the itinerary change in the same session is the natural next beat. - Per-type decay windows. A dietary restriction should probably never decay; a restaurant opinion should decay fast. Right now one window governs everything.
- Add Mastra's memory primitives (working memory, semantic recall) over Elasticsearch, so conversational continuity and the episodic record are two clearly separated layers.
- Learn the window instead of tuning it. The crossover point between a decision and its replacement is measurable from the data — it shouldn't be a hand-set constant.
Built With
- ai-agents
- claude
- elastic-cloud-serverless
- elasticsearch
- esql
- hybrid-search
- mastra
- node.js
- openrouter
- rag
- semantic-search
- typescript
- vector-search
Log in or sign up for Devpost to join the conversation.