Inspiration

Most "agent memory" demos fix forgetting. We went after the opposite failure: an agent that remembers everything, weights it all equally, and confidently acts on preferences you abandoned months ago. An agent that recalls nothing gives you a generic answer and you correct it. An agent with flat memory gives you a confidently personalized answer that's a year out of date — and you don't catch it, because it sounds like it knows you.

Here travel-agent project with a corpus of real landmark data, so we asked: what happens when the traveler changes and the memory doesn't?

What it does

A travel planner backed by eleven months of one traveler's decisions in Elasticsearch. Four of those decisions reversed:

Stale original Age Replacement Age
Walk the entire city, 20k steps/day 330d Knee injury — 3-mile cap, transit-adjacent 45d
Museum-first itineraries (SFMOMA) 300d Stop routing me through museums 25d
Pre-book Alcatraz and the Zoo 280d Skip ticketed attractions 35d
Late shows at The Fillmore 250d Nothing scheduled past 9pm 20d

Ask it to plan a free day in San Francisco. With flat memory it books an 8pm show and a 20,000-step walking route — for someone with a torn knee. With recency-weighted recall it returns a transit-adjacent, ticket-free, museum-free day that ends in bed by 9.

Nothing is deleted and nothing is edited. Supersession is a retrieval problem, not a storage problem.

How we built it

Mastra for the agent loop, with remember and recall as tools the model chooses to call — visible as explicit steps in the Studio trace rather than hidden context stuffing.

Elasticsearch as the memory store, with semantic_text computing embeddings server-side at index time. Recall is one ES|QL query:

  • FORK runs two branches in parallel — BM25 over title/content/tags, and semantic search over the embedded fields
  • FUSE LINEAR merges them at 0.3 BM25 / 0.7 semantic, because travel preference is prose rather than identifiers
  • DECAY(created_at, NOW(), 1080 hours) multiplies the fused relevance score by recency

Data: 20 real San Francisco landmarks from an existing travel-agent project, with a 15-entry episodic memory layer authored on top spanning 3 to 330 days.

Challenges we ran into

The dataset had to be built to make the wrong answer win on merit. Our first instinct was to write short corrections. That produced a demo where decay changed the ranking from slightly-right to more-right — undramatic and unconvincing. So we wrote the superseded entries ~3x longer than their replacements (749 chars vs 224), packed with specific rationale. Now BM25 and the embeddings both legitimately prefer the stale memory, and only time-weighting can overturn it.

Retrieval was correct while the answer was still wrong. The agent recommended a Twin Peaks hike to a user with a knee injury. The mobility memory ranked 6th; the tool's default limit was 5. Retrieval quality and retrieval quantity turned out to be separate failure modes — we raised the limit to 8.

The LLM kept rescuing bad retrieval, which broke our demo. With decay off, the agent often reasoned its way out of the contradiction on its own. That made the before/after non-deterministic and unusable as a live demo. We moved the demo down to the retrieval layer, which is pure ES|QL and reproducible to four decimal places.

Accomplishments that we're proud of

We measured our own failure rate instead of asserting a win. With decay off, the agent resolved three of the four reversals correctly by reasoning over the retrieved contradictions. It missed the fourth and booked a concert the night before a hackathon.

That number is the actual argument for this project: reasoning over stale memory works until it doesn't, and you can't tell which run you're getting. Decay at retrieval time makes correctness structural instead of lucky.

We can also show the window is a true half-life from our own data — the 45-day memory scores 0.7750 undecayed and 0.3899 at a 45-day window. Exactly half.

What we learned

  • Decay multiplies relevance rather than replacing it. Our 5-day trip context stayed ranked #1 in both conditions because it was recent and relevant. Decay is not recency bias.
  • DECAY's third argument must be a time_duration (1080 hours), not a date_period.
  • semantic_text on Serverless removes the entire embedding pipeline — no model choice, no vector generation, no client-side embedder.
  • Tuning is not just the knobs. The decay window, the fusion weights, the result limit, and the agent's conflict-resolution instructions are four separate levers, and the limit was the one that actually broke us.
  • Storing the age as a real timestamp rather than a staleness flag is what makes the result honest — there is nothing in the index marking a memory as superseded for the query to key off.

What's next for Superseded - episodic memory for travel planning

  • Write-path in the demo. The remember tool already works; capturing a reversal live and watching the itinerary change in the same session is the natural next beat.
  • Per-type decay windows. A dietary restriction should probably never decay; a restaurant opinion should decay fast. Right now one window governs everything.
  • Add Mastra's memory primitives (working memory, semantic recall) over Elasticsearch, so conversational continuity and the episodic record are two clearly separated layers.
  • Learn the window instead of tuning it. The crossover point between a decision and its replacement is measurable from the data — it shouldn't be a hand-set constant.

Built With

Share this project:

Updates

Submission history