Inspiration

Financial-crime decisions get audited long after they are made. A regulator can ask why a wire was flagged eight months ago, and the firm has to reconstruct exactly what the system knew at that instant: the alert, the context it recalled, the reasoning it applied.

Agents are terrible at this. The usual stack keeps embeddings in a vector database and state in Postgres, so the two drift apart. Rows get overwritten, history disappears, and there is no consistent point-in-time view across both. Worse, when the memory layer stalls, an agent does not degrade politely. It stops, or it acts twice and files a duplicate suspicious-activity report.

So we inverted the usual design. Instead of bolting memory onto an agent, we made the agent's entire memory an append-only ledger on CockroachDB, and built the product around what that makes possible.

What it does

Ledger is a fraud and AML triage agent whose every decision can be replayed exactly.

For each incoming alert it claims the case in a serializable transaction, writes the observation to the append-only ledger, recalls similar typologies and its own prior rulings through a distributed vector index, asks Amazon Bedrock for a flag / clear / escalate decision, then records that decision with the cluster logical timestamp it saw.

Because memory, vectors, and coordination all live in one strongly consistent store, the console can do things a split stack cannot:

  • Exact replay. Reconstruct the memory behind any past decision two independent ways, from the durable ledger and from CockroachDB's AS OF SYSTEM TIME, then compare them element by element on memory ID, kind, content, and timestamp. Matching row counts would not be proof.
  • Memory that compounds. On every alert the agent recalls its own prior decisions on semantically similar cases and cites them, so a repeated pattern gets consistent treatment. Those precedents are written into the ledger and reappear on replay.
  • Tamper-evident history. Each memory row extends a per-case hash chain that binds the previous hash, the IDs, the agent, the content, the embedding, and the metadata. One call re-walks every chain and proves nothing was altered, reordered, or deleted.
  • Counterfactual replay. Add a typology the agent never had, and ask whether the old decision would change today. The original snapshot stays untouched.
  • Semantic memory map. Server-side PCA over the stored VECTOR(256) embeddings, so you can see the clusters the agent actually recalls by.
  • Concurrency you can watch. Race twelve agents at one case, or launch a swarm across a batch. Exactly one agent ever wins a case, so no duplicate report is filed.
  • An auditor's read path. A compliance reviewer queries the same tables in natural language through the CockroachDB Cloud Managed MCP Server, with write tools disabled.

How we built it

CockroachDB is the memory layer, not a cache. known_patterns.embedding and agent_memory.embedding are VECTOR(256) columns with vector indexes, and recall runs embedding <-> $query similarity search. The same operator works under AS OF SYSTEM TIME, which is the trick that makes faithful replay possible: you can ask what the agent would have recalled at a past timestamp, vectors included, in one statement. case_claims uses serializable isolation with a retry loop so single-winner claiming is the database's guarantee rather than application glue.

Amazon Bedrock does the reasoning. Nova Pro decides through the Converse API with forced tool use, so the output is always structured and an unparseable response escalates instead of guessing. Titan Text Embeddings v2 produces the 256-dimensional vectors. Both calls are wrapped in exponential backoff with jitter for throttling.

AWS Lambda serves it. The FastAPI app runs as a container image through Mangum behind a public Function URL, with the image in Amazon ECR.

The auditor path is MCP. A service account connects an MCP client straight to https://cockroachlabs.cloud/mcp, scoped to the cluster. The agent writes the ledger; the auditor reads it back and can even walk the hash version of each row independently of the app.

Backing all of it: 53 offline tests over the hash chain, replay comparison, canonicalization, PCA, and API validation, plus three runnable proofs that execute against the real cluster for time-travel recall, serializable concurrency, and crash recovery.

Challenges we ran into

Our tamper-evidence was weaker than our claim. The first hash chain covered memory ID, kind, and content. That sounds fine until you notice the embedding drives recall and the metadata holds the cited precedents. Both could be rewritten while verification still reported "intact." We introduced a v2 hash that binds the embedding and metadata too, kept old v1 rows verifiable under their own rules, and made the integrity endpoint report coverage per version instead of implying the old rows were protected. Canonicalizing the vector was subtle: CockroachDB stores VECTOR as float32, so a written float64 comes back rounded and the recomputed hash misses unless both sides fold through the same precision.

Replay agreement was checking the wrong thing. It compared row counts. Two reconstructions can have identical length and different content, so it now compares positionally on every field and reports each mismatch.

Server-Sent Events do not stream through Lambda. Mangum buffers the response, so the live audit feed worked locally and silently died once deployed. The console now detects that within six seconds and falls back to polling the decisions ledger, labelling which transport is active.

TRUNCATE on a vector-indexed table is a background schema job. It once fired late and wiped a freshly seeded demo dataset. Resets are now DELETE-based in foreign-key order.

Role scoping for MCP is not obvious. With Cluster Developer, cluster metadata reads succeed while every SQL tool returns unauthorized, which looks like a broken connection rather than a permissions gap. Cluster Operator or Cluster Admin is required.

Small AWS accounts cannot reserve concurrency. Every account must retain ten unreserved executions, so on a quota of exactly ten the deploy script's safety cap was rejected. It now clamps or skips the reservation and treats the account quota as the ceiling.

Latency taught us to pool. Connecting per request from far outside the primary region cost about 1.1 seconds. A warm pool brought live database latency to roughly 2.5 milliseconds.

Accomplishments that we're proud of

The replay is honest. It does not assert that reconstruction worked, it compares two independent reconstructions field by field and shows you the result, including disagreement.

The integrity story survived our own audit. We found the gap between what we claimed and what the hash actually covered, fixed the hash, and wrote down precisely what older rows do and do not protect. Same with the MCP path: the managed server exposes write tools, so instead of repeating "read-only by default" we disabled them and documented how the constraint is enforced.

The concurrency guarantee is demonstrable in one click. Twelve agents, one winner, no duplicate report, enforced by serializable isolation rather than a lock we wrote.

And an outside auditor can verify the ledger without touching our application, which is the whole point of calling it a black box.

What we learned

Point-in-time queries change what a memory layer is for. Once AS OF SYSTEM TIME works on the same table as your vector index, "what did the agent know then" stops being an archival problem and turns into a single query. That single capability drove the replay, the counterfactual, and the audit feed.

We also learned to be suspicious of our own README. Several claims were true of the design and not of the code: hashing that skipped decision-driving fields, replay that counted rows, streaming that only worked locally, an MCP path called read-only when the server offered write tools. Every one was caught by asking what the code actually enforced. For anything auditable, the honest narrower claim is worth more than the broad one, because a reviewer can check.

Finally, the difference between a demo and a product here is failure behavior. Crash recovery, throttling backoff, cold-start warming, and a transport fallback are not polish. They are the reason the thing still works when you show it to someone.

What's next for mende

  • Move the write path behind real authentication and put the database credential in Secrets Manager with a KMS key. The public unauthenticated Function URL is a judging convenience, not a design.
  • Publish the hash chain's head per case to an external notary so integrity does not rest on the same system that stores the rows.
  • Use OAuth with a read-only scope for the auditor connection, so read-only is enforced by the authorization server instead of client configuration.
  • Extend replay past the garbage-collection window by materializing recall snapshots, so a decision from years ago replays with its vectors intact.
  • Multi-region agents claiming from the same ledger, to show that the single-winner guarantee holds across regions and not just across processes.
  • Feed analyst overrides back in as labelled memory, so the agent's precedents reflect what humans actually upheld.

Built With

  • agentic-memory
  • amazon-bedrock
  • amazon-ecr
  • amazon-nova
  • amazon-titan-embeddings
  • append-only-ledger
  • as-of-system-time
  • aws-lambda
  • cockroachdb
  • cockroachdb-cloud
  • distributed-vector-index
  • fastapi
  • hash-chain
  • javascript
  • mangum
  • model-context-protocol
  • numpy
  • pgvector
  • psycopg
  • python
  • serializable-transactions
  • sql
  • uv
  • uvicorn
  • vector-search
Share this project:

Updates

Submission history