Inspiration
I work at the SJSU IT servicedesk. Thirteen of us rotate through shifts, and we answer the same categories of calls all semester — eduroam drops, Duo lockouts, Canvas access. The problem isn't that we don't find fixes. It's that the fix one of us finds on Monday evaporates the moment they clock out. Whoever picks up the phone Thursday starts from zero, re-deriving something a teammate already solved three tickets ago.
And there's a sharper version of the problem that only shows up when you try to fix the first one: with 13 people logging fixes under call-volume pressure, a fix that worked once by luck looks identical to a fix verified ten times. If an AI agent repeats a lucky guess to all twelve other techs as settled procedure, it hasn't solved tribal knowledge — it has scaled the misinformation.
So Groundtruth's rule is in the name: memory has to earn the right to be believed.
What it does
Groundtruth is a Mastra agent whose long-term memory lives entirely in Elasticsearch:
- Recalls any teammate's fix with hybrid search — exact terms ("eduroam", "error 1132") through BM25, vague gists ("the WiFi thing that kicks students off") through server-side semantic embeddings — fused in one ES|QL query.
- Prefers recent truth over stale persuasion. Recall scores are decay-weighted, so a terse two-week-old "the auth backend migrated, old fix superseded" outranks a beautifully-documented procedure from June that is now wrong.
- Refuses to oversell a guess. New memories are born
unconfirmedand score-penalized (×0.5) until a different tech verifies the fix in an independent case — the same tech confirming themselves is rejected. Watch it promote a lead to confirmed knowledge live, with both names on it. - Knows when something is systemic. A
count_signaturetool runs a live ES|QLSTATSaggregation over the ticket log: "this signature hit 11 times this week across 4 techs — 5.5× last week, escalate it." Computed, not remembered. - Knows who it's talking to. Mastra's working memory + semantic recall (vectors on Elasticsearch, embeddings computed by the cluster) keep each tech's profile across sessions, so fix attribution is automatic.
How I built it
The retrieval core is one ES|QL pipeline: FORK (BM25 branch + semantic branch) → FUSE LINEAR → DECAY on recency → a trust CASE multiplier — extended from the starter's remember/recall pattern with the trust fields, a confirm_fix tool, and the aggregation tool. Embeddings are semantic_text, computed server-side by Elasticsearch Serverless, so there is no embedding pipeline at all. The executed ES|QL is returned inside every tool output, so it's visible in the Mastra Studio trace — the retrieval isn't a black box, it's on screen.
Challenges
- Tuning was real, not decorative. With default RRF fusion, cross-topic semantic "vibes" matches crowded the query's own older answer out of the results entirely — the agent couldn't even see the fix it was superseding. Shifting to
FUSE LINEARwith BM25 at 0.45 fixed it, because servicedesk language is ID-heavy. - The starter's recall only returned memory titles. The agent honestly told me "a fix exists but the memory is thin on procedure" — it literally couldn't see the content field. Good honesty, bad answer; fixed by carrying content through the ES|QL
KEEP. - Aggregation precision. An OR-match on "eduroam auth loop" blended three different issues into one scary number. AND-operator matching plus an English analyzer (so "loop" matches "loops") made the counts exact.
What I learned
Agent memory is a ranking problem, not a storage problem. Storing everything is trivial; deciding that a 4×-confirmed fix should lose to a 2×-confirmed one because the world changed on Aug 7 — that's where Elasticsearch's scoring machinery (decay, fusion weights, custom multipliers) does work an LLM can't fake.
What's next
Per-tech document-level security (private notes vs. shared catalog), ingesting our real ticket export, and a supersession edge: letting a confirmed memory formally retire its predecessor with an audit trail.
Built With
- claude
- elastic-cloud
- elasticsearch
- esql
- libsql
- mastra
- node.js
- openrouter
- typescript
- zod
Log in or sign up for Devpost to join the conversation.