Inspiration
Most multi-agent demos treat "memory" as a chat-history footnote - a buffer that resets the moment the process restarts, in a domain where nobody's livelihood depends on it remembering correctly. Market surveillance is the opposite kind of domain: an investigation team needs to know whether a pattern has been seen before, across every case anyone has ever worked, not just the current conversation - and a regulator needs an immutable, queryable record of exactly what was asked and answered, indefinitely.
We started from AWS's own reference architecture, and asked a deliberately provocative question for the CockroachDB × AWS Hackathon: what if the memory layer wasn't a bolt-on cache sitting next to the real system of record, but was the system of record - one database, under transactional guarantees, holding conversation state, audit history, and semantic case memory all at once? And what if that database was one built to survive the kind of infrastructure failure a compliance system can't afford to go down during?
What it does
You ask a market-surveillance question - "What was the trading activity for
AAPL on March 15, 2024, and which brokers were most active?" - and an
orchestrator built on LangGraph routes it to the right specialist agents out
of eight: security_monitor, broker_monitor, risk_monitor,
intel_analyst, compliance_officer, case_triage, audit_reviewer, and
memory_ops. Each one reasons through a Strands Agent calling Anthropic's
API directly.
The part that makes this more than a chatbot: before the orchestrator ever
sees your question, the system asks its own memory whether it's seen
something like it before - a real semantic similarity search over every
past investigation, using CockroachDB's native VECTOR type and its
distributed C-SPANN index. If it finds something, that context shapes how
the question gets routed and answered. When the investigation finishes, the
finding is embedded and written back, so the system's memory compounds with
every use instead of resetting per session.
Two agents put that memory directly in their own hands, as a tool call
rather than something the graph does automatically: case_triage explicitly
queries case memory to decide how urgent the current investigation is based
on how similar past ones resolved, and memory_ops is wired to the
CockroachDB Cloud Managed MCP Server so it can answer "how many cases do
we have on file for this broker?" by actually querying the live cluster,
read-only.
Everything is checkpointed after every graph node (a crash or an
AgentCore replica restart mid-investigation resumes exactly where it
stopped) and every completed turn is written to a separate, human-readable
audit table a compliance reviewer could SELECT * FROM directly.
A modern React UI makes all of this visible instead of an implementation detail: a streaming Chat view with a live agent-pipeline graph and explainable citations, and a Dashboard with a Vector Memory Map - every case embedding projected to 2D and searchable live - and a Time Travel scrubber over a session's actual LangGraph checkpoint history. That UI is reachable two ways: run it locally with zero AWS involved, or hit a live, publicly-hosted copy on Amazon ECS - no installation required to try it.
How we built it
The market-surveillance domain - agent personas, mock report catalog,
report schemas - is forked from this repository's existing
sample-langgraph-strands-market-analysis sample. Everything about memory,
model inference, and the deployment automation around both was built
fresh for this hackathon:
- LangGraph orchestrates the whole thing as a
StateGraph: arecall_case_memorynode runs before the orchestrator, specialist agents fan out and back in, a synthesizer writes the final answer, and apersist_case_memorynode closes the loop. - Strands Agents do the reasoning, each backed by an
AnthropicModelpointed straight at Anthropic's API - Bedrock is never in the inference path. - CockroachDB is the only system of record, doing three distinct jobs:
AsyncCockroachDBSaverfor LangGraph's own checkpoint state,CockroachDBChatMessageHistoryfor the durable audit log, andAsyncCockroachDBVectorStore(backed by a C-SPANN distributed vector index) for long-term semantic case memory. - Amazon Bedrock AgentCore hosts the agent runtime itself:
api.pywraps the LangGraph workflow in aBedrockAgentCoreAppentrypoint. Around it, a small ring of supporting AWS services handles build and hosting, never inference: AWS CodeBuild builds the ARM64 container natively - no local Docker or QEMU needed - straight from this GitHub repo, pushes it to Amazon ECR, and provisions the AgentCore runtime viaboto3. The whole path is one script (deployment/deploy-codebuild.sh) that also runs unattended from a stock AWS CloudShell session (deployment/cloudshell_deploy.sh) for a genuinely one-shot deploy with nothing installed locally. - A FastAPI layer sits between the React UI and the agents, because a
browser can't safely hold AWS credentials to call AgentCore's
invoke_agent_runtimedirectly. It gives the UI one streaming chat contract regardless of whether the backend is running the workflow in-process (local dev, zero AWS) or proxying to a deployed AgentCore runtime - same UI, same API, no code changes to switch. - The same FastAPI + React UI also has its own, independent deployment path:
Amazon ECS Express Mode stands up a Fargate service behind an
internet-facing Application Load Balancer, with automatic HTTPS and
autoscaling, from a single
create-express-gateway-servicecall. We originally scoped this as an AWS App Runner deployment, but App Runner is closed to new customers as of 2025 - ECS Express Mode is the direct replacement AWS itself documents, and it's whatmake deploy-webactually drives. - The ccloud CLI provisions the CockroachDB Cloud cluster and SQL user end to end from the terminal, with JSON output at every step.
- The React + Vite + TypeScript + Tailwind UI's charts are hand-rolled (no charting library) against a categorical palette validated for a dark surface and colorblind accessibility.
- Project documentation is a MkDocs Material site, built and published to GitHub Pages automatically by a GitHub Actions workflow on every push to the public repo.
Challenges we ran into
- A silently broken recall path.
AsyncCockroachDBVectorStore'sasimilarity_search_with_relevance_scores()raisesNotImplementedErrorin the installedlangchain-cockroachdbversion - and our own code had it wrapped in a broadtry/exceptthat logged a warning and returned an empty list. Recall was silently failing shut on every single call, with nothing anywhere in the app surfacing that it was broken - the exact failure mode you'd never catch from a demo that "looked fine." We found it while double-checking the project's own central claim before writing this submission, and fixed it by switching toasimilarity_search_with_score()(raw cosine distance, since the store usesDistanceStrategy.COSINE) and converting to a similarity score ourselves. - Distributed vector indexing needs a cluster setting turned on first -
CREATE VECTOR INDEXfails outright untilfeature.vector_index.enabledis set, which isn't obvious from the CREATE INDEX syntax alone. - A real API-drift bug:
ainit_vectorstore_table()was being called withdistance_strategy=andoverwrite_existing=keyword arguments that don't exist on the installed library's actual signature - aTypeErroron every single startup until we read the installed source directly to find the real parameter list. - Windows compatibility: psycopg's async driver refuses to run under
Windows' default
ProactorEventLoopoutright. Fixed with a small, explicitWindowsSelectorEventLoopPolicyswitch applied as early as possible in every entrypoint - the FastAPI server and every standalone script. - A nested-event-loop bug in our own schema-init script: a synchronous
convenience wrapper in the chat-history library calls
asyncio.run()internally, which raises when called from code that's already inside a running event loop. Fixed by awaiting the library's async implementation directly instead of the sync wrapper. - Keeping AWS honestly out of the inference path. It's easy to say
"Bedrock isn't used for inference" - we made sure it was actually true by
checking that the deployment IAM policy carries no
bedrock:InvokeModelpermission at all, so the architecture can't quietly regress to using Bedrock without a visible permissions change. - IAM eventual consistency broke automated deploys, twice, in two
different services. A freshly created CodeBuild service role's inline
policy hadn't propagated by the time the first build kicked off, producing
an opaque
CLIENT_ERROR/ACCESS_DENIEDwith no CloudWatch log group to even inspect. We only found the real cause by querying the build's rawphases[*].contextsinstead of trusting the failed-phase filter, which came back empty. Fixed with a longer post-creation sleep and a retry loop that specifically recognizes the propagation-error message pattern and retries instead of failing outright - the exact same race showed up again standing up ECS Express Mode's IAM roles, and got the same fix. - A first-time ECS account doesn't have its service-linked role yet.
create-express-gateway-servicefailed with "Unable to assume the service linked role" the first time we ever used ECS Express Mode in this accountAWSServiceRoleForECSisn't guaranteed to already exist, so the deploy script now checks for it and creates it explicitly before proceeding.
- A judge's password with special characters got silently corrupted by
our own deploy script. Loading
.envwith a plain bashsourceruns it as a shell script, so aJUDGE_ACCESS_PASSWORDcontaining shell metacharacters got mangled before it ever reached the deployed container - the login screen was rejecting a password that had already been altered upstream, not the one actually typed. Fixed by parsing.envwithpython-dotenv'sdotenv_values()instead of sourcing it, the same safe approach the app itself already used. - CockroachDB Cloud's
sslmode=verify-fullneeds a CA cert that simply doesn't exist in any container we ship. libpq's default lookup path (~/.postgresql/root.crt) only ever gets populated by a developer manually downloading it from the Cloud console - never by an automated CodeBuild pipeline. We layered three fixes, each catching what the last one missed:sslrootcert=system(works when the cluster's cert chains to a publicly-trusted CA), an optionalCOCKROACHDB_CLUSTER_IDenv var that downloads the cluster's actual CA cert at container startup, and the same cert baked into the image at build time via acurlstep in both Dockerfiles (so there's zero runtime network dependency at all). Even with the correct cert in hand, we still hitSSL error: certificate verify failed- which turned out to be the cluster's TLS proxy not sending its intermediate certificate during the handshake, a server-side chain-completeness issue no client-side cert file can fix. The pragmatic unblock:sslmode=requirefor this deployment - still fully encrypted, just without chain verification against a SaaS endpoint whose host/IP range is already known.
Accomplishments that we're proud of
- CockroachDB as the only memory system in the stack - no Redis, no separate vector database, no AgentCore Memory. Three genuinely different jobs (workflow checkpointing, audit history, semantic recall) done by one database under one set of transactional guarantees.
- Finding and fixing a real, ship-blocking bug in the project's central feature before treating the submission as done, instead of demoing a path that silently did nothing.
- Turning "introspect your own memory cluster" into an actual tool call -
the
memory_opsagent answers questions about its own infrastructure by querying it live through the CockroachDB Cloud Managed MCP Server, not by guessing. - A UI that makes a distributed vector index something you can look at - the Vector Memory Map is a real, live, searchable 2D projection of the actual embeddings backing recall, not a static illustration.
- One codebase, two deployment modes, zero behavior difference:
make devruns the exact same LangGraph workflow with zero AWS involved, andmake deployputs the identical workflow behind AgentCore - the UI can't tell the difference. - A fully automated, one-command deploy path - from a stock AWS CloudShell session with nothing installed locally, to a live AgentCore runtime and a publicly reachable web UI on Amazon ECS - hardened against three real IAM/ECS bootstrap races we hit while actually running it, not just the happy path.
What we learned
- A
try/exceptaround an async library call that only logs a warning is exactly how a broken feature ships looking like a working one - test the actual return value against the exact installed library version, not just that the call doesn't raise in your dev loop's happy path. - Running vector search on the same OLTP database that holds your operational data is a genuinely different value proposition from bolting on a dedicated vector database: the fact "this investigation happened" and the fact "this investigation is now recallable" become transactionally consistent for free.
- Async Postgres-wire drivers and Windows' default event loop don't mix by default, and you won't discover that from a Linux development machine - it's worth testing (or at least defensively coding for) cross-platform from day one.
- Least-privilege IAM is worth enforcing as a literal grep of the policy file at review time, not just an architectural intention that's easy to quietly violate later.
- Automation that touches freshly created IAM roles needs to treat eventual consistency as the normal case, not an edge case - a longer sleep plus a targeted retry on the specific propagation-error pattern is cheap insurance against a race that's otherwise nearly impossible to reproduce on demand, and we hit it twice, in two unrelated services.
- Never
sourcean untrusted.envfile as a shell script - parse it instead. It's exactly the kind of bug that only shows up when a real value has special characters in it, which is exactly what a generated judge password does. - "The CA cert is correct" and "the connection verifies" are two different
claims - a valid root certificate doesn't help if the server never sends
the intermediate that bridges its leaf cert up to that root. When
verify-fullfails even with a known-good cert in hand, the next question is what the server is presenting, not which file the client is holding -openssl s_client -showcertsanswers that in one command, before reaching for another client-side workaround.
What's next for Market Memory Mesh Agents
- Benchmark C-SPANN recall latency at real scale (thousands of cases, not a handful of seeded demos) and tune the index's build parameters accordingly.
- Give
case_triageandcompliance_officerreal write-capable tools behind a human-in-the-loop approval step - compliance actions need sign-off, not just a recommendation. - Explore multi-region CockroachDB placement for the audit trail itself, so the compliance record survives losing an entire cloud region, not just a single node.
- Replace the mock report catalog with real market-data connectors.
- Broaden the MCP toolset beyond CockroachDB's own MCP server to other compliance-relevant, read-only data sources.
Built With
python
typescript
react
vite
tailwindcss
fastapi
langgraph
langchain
strands-agents-sdk
anthropic-claude
claude
cockroachdb
cockroachdb-cloud
distributed-sql
postgresql
vector-search
psycopg
fastembed
boto3
aws
amazon-bedrock-agentcore
aws-codebuild
amazon-ecs
amazon-ecr
docker
mcp
mkdocs
github-actions
github-pages
Grouped, for reference:
- Languages: Python, TypeScript
- Frontend: React, Vite, Tailwind CSS
- Agent orchestration: LangGraph, LangChain, Strands Agents SDK
- Model inference: Anthropic Claude (via the Anthropic API, called directly - no Bedrock model invocation)
- Backend: FastAPI
- Database: CockroachDB / CockroachDB Cloud (checkpointing, audit log,
and C-SPANN distributed vector index for semantic case memory), accessed
via
langchain-cockroachdb+psycopg - Embeddings: fastembed (local ONNX, no external embedding API)
- Cloud/hosting: AWS (
boto3), Amazon Bedrock AgentCore (agent runtime hosting), AWS CodeBuild (native ARM64 container builds, no local Docker), Amazon ECR, Amazon ECS Express Mode (public web UI hosting), IAM, AWS CloudShell (one-shot unattended deploy path) - Protocols: Model Context Protocol (MCP) - the CockroachDB Cloud
Managed MCP Server, wired into the
memory_opsagent - Containerization: Docker (ARM64 build for AgentCore, x86_64 build for the web UI)
- Docs: MkDocs Material, published to GitHub Pages via GitHub Actions https://akashtalole.github.io/MemoryMesh-Agents/
Built With
- amazon-bedrock-agentcore
- amazon-ecr
- amazon-web-services
- anthropic
- boto3
- claude
- cockroachdb
- cockroachdb-cloud
- ecr
- ecs
- fastapi
- fastembed
- langchain
- langgraph
- mcp
- python
- react
- strands
- strands-agents-sdk
- tailwindcss
- typescript
Log in or sign up for Devpost to join the conversation.