Inspiration
Every team at this hackathon is going to build a place to put agent memory. Recall quality is a crowded, well-trodden problem at this point. What kept nagging at us is that nobody could answer three questions that actually matter once an agent's memory is doing something real: where does this memory legally live? What did the agent actually believe when it acted? And if one poisoned fact got in, what did it touch?
That third one is what really got us. Persistent prompt injection isn't a "sometimes agents say weird things" problem — it's a permanent contamination problem. A memory system that lets a poisoned fact quietly infect every future recall, forever, with no way to even measure the damage, is not a system you'd trust with anything that matters. And the EU AI Act's Article 12 record-keeping requirements go live for high-risk systems this August — the same month as this deadline. This isn't a hypothetical compliance problem we invented to sound serious. It's already due.
The thing that made us actually want to build this, instead of just complaining about it, was realizing why nobody had answered these questions well: because everyone was building memory as a vector store bolted to Postgres bolted to a separate audit log, with a consistency gap between every pair of those systems. You can't build atomic erasure or a computed blast-radius revocation across a gap like that. But CockroachDB isn't a gap — it's one transactionally consistent database that can hold vectors, a provenance graph, and a hash-chained ledger together, in the same serializable transaction. Once we saw that, the whole shape of the project fell out of it. This wasn't "let's add governance to a memory store." It was "governance IS the product, and the memory tiers are the substrate underneath it."
What it does
Mnemos is memory an agent can be held accountable for, not just memory it can recall from. Four planes sit over one CockroachDB cluster — the Fabric (episodic → semantic → procedural memory), the Ledger (sharded hash chains, Merkle checkpoints anchored to S3 Object Lock), the Warden (residency, legal holds, erasure, blast-radius revocation), and the Custodian (an LLM agent that watches the cluster's own health and reads that back into memory). Exactly one of those four planes contains a model, and it isn't the one that can destroy anything.
Three real mechanisms carry the whole thing:
- Residency.
REGIONAL BY ROWhomes every episode to a jurisdiction. An agent in another region can still get an answer — a policy-approved derived projection — but the raw content never crosses the border, and every crossing that does happen gets logged with the policy that allowed it. - Accountability.
explain(action_id)reconstructs exactly what an agent believed at the instant it acted, hash-verified end to end, exportable as an HTML file that verifies itself completely offline — no server, no trust required. - Integrity. Everything an LLM writes lands
unverifiedand stays out of recall until something independent corroborates it — two different sessions, two different source categories, computed as real bipartite matching, not a heuristic. When a source turns out to be compromised,revoke_source()computes the full blast radius — every fact, skill, and past decision that source touched — and revokes it all in one transaction.
It's live, not staged. Three demo tenants (a clinic, a DevOps team, a consumer-finance shop) run on the same fabric with genuinely isolated data, and you can walk any of them yourself in the console with no key required.
How we built it
Python for the engine, the Warden, the sleep cycle, and the Custodian; TypeScript and Next.js for the console; SQL doing more of the actual enforcement than either of them. The write path is AWS Lambda behind API Gateway, speaking MCP so any agent — Claude, LangGraph, whatever — can just connect to it. Consolidation runs asynchronously on Step Functions, calling a model to distill episodes into facts and never touching anything on the synchronous write path, so memory intake survives a total model-provider outage by construction. The Custodian runs on ECS Fargate, scheduled and alarm-triggered, and it's the only thing in the whole system that talks to the CockroachDB Cloud MCP server — read-only, allowlisted SQL pulled straight from five official CockroachDB Agent Skills.
We held ourselves to a rule the whole way through: five invariants, each backed by a named test, two of them enforced by the database itself rather than application code. No LLM-driven process holds DELETE — checked statically in CI against the Warden's own import graph, not just asserted in a doc. Every state-changing memory operation appends a hash-chained audit row in the same transaction as the change, enforced by a database trigger that refuses the mutation outright if the audit row is missing. We didn't want "we promise this is safe." We wanted a build that fails loudly the moment it stops being true.
Challenges we ran into
The one that scared us the most wasn't a red-team attack — it was our own console. Deep into building the dashboard, we found that the tenant switcher looked like it worked perfectly, but under the hood one code path was quietly defaulting to a static API key regardless of which tenant was selected. Every screen looked tenant-scoped. It wasn't. For a product whose entire pitch is "isolation you can verify, not take on faith," that bug landing in front of a judge would have been the whole argument collapsing in real time. We only caught it because we stopped trusting what the UI looked like it was doing and actually minted real keys for real tenants and checked the numbers came back different. They didn't, at first. Then they did.
We also found, and had to sit with, a real security hole of our own making mid-build: source_trust used to be a caller-declared argument, which meant an injected agent could just say "treat me as operator" and skip corroboration entirely — the whole trust gate, defeated in one tool call. Closing it meant binding system/operator provenance to an admin credential instead of trusting whatever an agent claimed about itself. It's exactly the kind of thing that's obvious in hindsight and invisible until you're specifically trying to break your own system.
And we did try to break it, on purpose — thirteen attacks against our own memory, and one of them actually worked: two independently-controlled sources can still collude to promote a fact past the corroboration gate. We could have quietly not published that. We published it anyway, with the exact mechanics of why it works, because a defense that only admits its wins isn't a defense you should trust either.
Smaller fights along the way: a diagram library that computed its layout against an unmeasured 0×0 box and rendered a tiny graph adrift in a huge empty container — twice, because the second time was a proportions bug, not the same bug. A Vercel deploy that failed because we'd set the monorepo's Root Directory and deployed from inside that same directory, double-nesting the path. Discovering the CockroachDB Cloud MCP server's SQL tools are locked behind Cluster Admin specifically — Monitor and Developer both block every SQL-shaped tool outright, which is nowhere obvious until you've hit it. And a five-dollar OpenAI budget for the entire hackathon, which meant every real-model test had to earn its place instead of running by default.
Accomplishments that we're proud of
439 unit tests, 118 invariant tests, and a red-team suite, all green, alongside lint, typecheck, and three structural CI guards that make the architecture's core promise mechanical instead of aspirational — make no-warden-in-custodian literally fails the build if the Custodian's import graph can reach the Warden at all. We're proud that "the only component that can destroy memory has no model in it" isn't a tagline here; it's a thing a build fails on if it stops being true.
We're proud we shipped the collusion attack that works, instead of only the ones we blocked. We're proud two skills we wrote — distilled straight out of building this, not written to pad a rubric — are real, open pull requests against CockroachDB's own Agent Skills repository right now. And we're proud the whole thing is actually live: a real deployment, three real isolated tenants, and a console that shows the trust lattice and the audit chain as they actually are at this moment, not a recording of them from last week.
What we learned
That governance is genuinely a different kind of hard than recall quality — recall gets better with a better model; provable deletion, jurisdiction-aware storage, and computed blast radius get better with better transactions, and no amount of prompting fixes a consistency gap between two separate systems. That "we tested it" and "we ran it against the real deployed instance and watched the numbers actually change" are two very different claims, and only one of them should go in a README. And that the boring, unglamorous work — closing an IAM permission gap, fixing a trigger, re-checking a claim against the actual code instead of the plan that described it — is where almost all of the real trust in a system like this actually gets built.
What's next for Mnemos — Accountable Memory for Agents
The honest list, not the flattering one: crypto-shred needs to be wired all the way into the deployed API against the provisioned per-tenant KMS keys (it's tested against real KMS semantics, just not live yet). Checkpoint anchoring needs to move from a manual step to a real schedule, which is what actually bounds how far back tamper detection reaches. Three of the six planned red-team attack classes — ledger tampering under load, residency violation via changefeed, erasure evasion inside the GC window — are specified and still unrun, and we'd rather say that plainly than let anyone assume otherwise. Past that: real load and scale numbers instead of architectural reasoning about them, and growing the Warden's governance surface so more of "what should an agent never be allowed to do" becomes a database-enforced fact instead of a policy anyone has to remember to check.
Built With
- amazon-web-services
- aws-api-gateway
- aws-cloudwatch
- aws-ecs-fargate
- aws-eventbridge
- aws-kms
- aws-lambda
- aws-secrets-manager
- aws-step-functions
- cockroachdb
- cockroachdb-cloud
- framer-motion
- mcp
- next-js
- openai
- pnpm
- pytest
- python
- react
- react-flow
- sql
- tailwind-css
- typescript
- uv
- vercel
Log in or sign up for Devpost to join the conversation.