Inspiration
Here is the failure Trace was built around. A team learns something the hard way. Maybe it came from an outage. Maybe a security review. Maybe three engineers spent an afternoon arguing through a trade-off until they finally understood why the obvious solution was dangerous. They make the right decision. They merge the fix. The reasoning ends up in a pull request. Then everyone moves on. Months later, another developer opens a completely reasonable PR that brings the same problem back. Not because they are careless. Because the code remembers what changed, but not why. That is exactly what we demonstrate with Trace. PR #4 establishes a security invariant:
Authorization caches must be invalidated immediately when access is revoked. A TTL alone is not enough. Trace captures that decision as
TRACE-MEMORY-00401, including its source, rationale, scope, confidence, security relevance, and embedding. Then PR #5 introduces a ten-minute permission cache. If you only read PR #5, it looks like a sensible performance optimization. If you remember PR #4, it is the same stale-authorization failure coming back in a different form. Trace remembers. It retrieves the old decision, recognizes the semantic conflict, cites the original PR, and blocks the review. That became the idea behind the whole project: Most AI coding tools help teams write code faster. Trace prevents teams from writing the wrong code again. ADRs and wikis can preserve decisions, but only when someone remembers to write and maintain them. Trace learns from work the team is already doing.Your codebase remembers why.
What it does
Trace is institutional memory for software teams. It follows a change from issue to production and gets more useful after every merge.
| Your team's situation | What Trace does |
|---|---|
| Developer creates an issue | Specforge finds relevant past decisions and failures, runs a pre-mortem, asks hard questions, and records implementation promises |
| Pull request opens | Guardkeeper retrieves governing memories and checks the diff for semantic conflicts, broken promises, security regressions, and architectural drift |
| New code recreates an old mistake | Trace cites the exact memory, rationale, and original source instead of giving a generic warning |
| A developer intentionally changes a decision | Reply Handler supersedes the old memory without erasing the history behind it |
| Pull request merges | Tracekeeper extracts durable decisions from the work and stores provenance, scope, confidence, relationships, and embeddings |
| Team wants a health view | Tracecast builds decision health, security inventory, dependency context, and onboarding signals |
| Someone asks "why is this built this way?" | Trace Ask answers from governed memory with specific decisions, sources, dates, and relationships |
The important part is that a Trace memory is not just a chunk of text in a vector store. It has a lifecycle. It has provenance. It has scope. It has confidence. It can be security-relevant. It can depend on, contradict, or supersede another memory. And it can actually change what the agent does. That last part became important enough that every Guardkeeper review produces a Memory Consequence Receipt. For our PR #5 example:
Memory changed this review: yes Governing memory:
TRACE-MEMORY-00401Memory-conflict findings: 1 Independent findings: 0 Counterfactual: without the selected institutional memory, the conflict finding would be absent. Retrieving something is not enough.Trace proves when remembering changed the review.
How we built it
Trace runs a real production-style path: GitHub → CockroachDB → Amazon SQS → Amazon Bedrock → GitHub
CockroachDB is the memory system
CockroachDB is not just a vector sidecar in Trace. It is the canonical source of truth. It stores:
- governed memories
VECTOR(1024)Titan embeddings- source provenance
- file and repository scopes
- memory relationships
- lifecycle state
- retrieval events and candidates
- agent tasks
- audit events
- feedback
- transactional outbox state
This matters because similarity and authority are not the same thing. A vector result can say: "this looks related." Trace also needs to know: "this memory is ACTIVE, belongs to this repository, applies to this file, came from this source, has this confidence, and has not been superseded."
Retrieval is hybrid, not just nearest-neighbor search
Amazon Titan Text Embeddings V2 creates 1024-dimensional embeddings. CockroachDB Distributed Vector Indexing retrieves the nearest tenant-scoped candidates. Then Trace scores them using:
- semantic similarity
- changed-file scope
- confidence
- security relevance
- previous feedback The strongest candidates go to Amazon Bedrock for semantic reasoning. Nova Pro is the primary reasoning model, with Mistral Large as the secondary model. The model is not allowed to invent database truth. It may select only memory IDs that actually came from CockroachDB. If Bedrock returns an unknown ID, Trace rejects the result. Every model response is schema-validated with Pydantic before it can affect stored state or review behavior. Repository content is treated as untrusted evidence, not instructions. ### Memory evolves instead of being overwritten Engineering decisions change.
So Trace does not pretend an old rule should remain true forever.
When a team intentionally replaces a decision, the new memory becomes active and the old one becomes SUPERSEDED.
That transition happens atomically in CockroachDB, and a SUPERSEDES relationship connects the new reasoning to the old reasoning.
You can still see what the team believed before and why it changed.
The runtime is built for retries and crashes
GitHub webhooks are verified with HMAC-SHA256 before Trace accepts them. Trace also checks repository identity, payload size, bot events, and delivery IDs. The task and its transactional outbox event are committed together in CockroachDB. Amazon SQS FIFO handles durable delivery. The runtime uses:
- repository-scoped idempotency
- FIFO deduplication
- database-backed task leases
- bounded retries
- visibility timeout changes
- encrypted dead-letter queues
- task checkpoints
- idempotent GitHub publication
If GitHub sends the same event twice, Trace does not create two effective tasks. If a worker posts a GitHub review and crashes immediately afterward, Trace finds the existing task marker instead of posting the review again.
We built the demo to be live, not theatrical
The public AWS Lambda judge console executes a fresh read-only path every time Run Trace is clicked: PR diff → Titan embedding → CockroachDB retrieval → Bedrock reasoning → consequence receipt It reports real stage timings, the selected memory, actual model ID, verdict, and counterfactual. It uses the same Guardkeeper path as the production runtime. It has zero write routes. There is no replay fallback. If a dependency fails, Trace shows: NO RESULT FABRICATED We would rather show a real failure than a fake success.
Independent verification with Managed MCP
We also built Trace Auditor using the CockroachDB Cloud Managed MCP Server. It has read-only access. Given a memory or retrieval ID, it can independently verify: memory → source → retrieval → selected candidate → task → published action So the application is not the only thing saying, "trust me, memory mattered."
The underlying records can be inspected directly.
Challenges we ran into
Similarity is not authority — the nearest vector is not always the decision that should govern the code. A slightly less similar security constraint may matter far more. That pushed us toward hybrid ranking with scope, confidence, security relevance, feedback, and lifecycle state.
Memory has to know how to become outdated — silently editing old decisions destroys history. We built explicit
ACTIVE → SUPERSEDEDtransitions instead.Retrieval does not prove consequence — an agent can retrieve the perfect memory and still produce the same answer it would have produced without it. That problem led to the Memory Consequence Receipt.
Models can sound confident while being wrong — Bedrock outputs are schema-validated, prompt versions and model IDs are recorded, and memory selection is constrained to IDs that actually came from CockroachDB.
GitHub delivery is not exactly-once — webhook duplication, transaction retries, worker crashes, and external GitHub side effects create ugly failure windows. Transactional outbox state, FIFO deduplication, leases, checkpoints, and idempotent publication handle them.
A convincing demo is easy to fake accidentally — replaying a successful result would have been simpler and more reliable during judging. We deliberately chose a fresh live path instead.
Memory has to be tenant-safe — organization and repository scope are applied before vector retrieval, not as an afterthought.
Accomplishments that we're proud of
- A governed memory model with provenance, confidence, scope, security relevance, relationships, feedback, and lifecycle
- Semantic conflict detection that catches the same failure pattern even when the implementation is worded differently
- A Memory Consequence Receipt that proves when institutional memory actually changed a review
- A real PR #4 →
TRACE-MEMORY-00401→ PR #5 memory-to-action chain - CockroachDB Distributed Vector Indexing tested against 10,000 production-schema 1024-dimensional rows
- Tenant-prefixed vector search with no full table scan in the production query proof
- CockroachDB Cloud Managed MCP used as an independent read-only audit surface
- Schema-constrained Amazon Bedrock reasoning with candidate-ID validation
- A real GitHub → CockroachDB → SQS → Bedrock → GitHub runtime
- HMAC verification, durable task admission, FIFO deduplication, bounded retries, DLQ handling, and least-privilege database roles
- A public AWS Lambda demo with no replay fallback and zero write routes
- 51 passing tests
- CI that runs tests and static checks, validates migrations and CloudFormation, builds the Python package and Docker image, and audits dependencies But the thing we are most proud of is much smaller than that list: A past decision changed a future review. That is what we wanted memory to mean.
What we learned
- Retrieval is not the product. Consequence is. Getting the right memory into context only matters if it changes what happens next.
- Organisational memory needs authority, not just similarity. Scope, confidence, lifecycle, provenance, and security relevance matter.
- Forgetting correctly is part of remembering. A healthy engineering organization changes its mind sometimes. The history should survive.
- Provenance changes how developers react to AI. "This looks dangerous" is easy to ignore. "This conflicts with a decision your team made in PR #4; here is the source and rationale" starts a useful conversation.
- Reliable AI is mostly boundaries. Transactions, schemas, idempotency, permissions, audit trails, and failure semantics are what make model reasoning usable in a real workflow.
- Keeping semantic and operational truth together simplifies everything. CockroachDB lets the vector, lifecycle, provenance, retrieval evidence, and action history live in the same transactional system.
- The best memory is slightly opinionated. A useful codebase should sometimes be able to say: "I know this looks reasonable. We already learned why it isn't."
What's next for Trace
- Cross-repository institutional memory with explicit scope boundaries
- Organization-level memory and policy controls
- IDE integration so governing decisions appear while code is being written
- Better retrieval-quality evaluation sets
- Richer dependency graphs between decisions
- Confidence and feedback evolution over time
- Operator-approved dead-letter replay
- Multi-region workers
- Deeper onboarding and Trace Ask experiences
- Measuring not only whether the correct memory was retrieved, but whether it actually prevented a repeated mistake The long-term goal is not to build another chatbot that happens to know your repository. We want Trace to feel like the engineer who has been on the team for years. The one who remembers the outage. The one who remembers the rejected shortcut. The one who can say:
"We tried this before. Here is what happened. Here is what we decided. And this PR is bringing the same problem back."
Except that knowledge should belong to the codebase, not to whoever still happens to work there.
Your codebase remembers why.
Log in or sign up for Devpost to join the conversation.