-
-
The incident board: live triage across services, each incident carrying its own state machine and agent reasoning timeline.
-
Same incident, same loop, twice with search_memory genuinely withheld from the control arm.Memory cites the precedent;control cites nothing.
-
EXPLAIN proves it: vector search on memories_embedding_cos_idx, no full table scan. The check a skeptical judge should ask for.
-
Killed after step 2, resumed by a fresh holder reading only CockroachDB checkpoints. Zero duplicate side-effects, no replayed model calls.
-
New incident recalls a 9-day-old precedent at 97% similarity in 2.5ms and proposes the fix that actually worked. Real run, real CockroachDB.
-
No precedent above threshold? It says so instead of inventing a match. The honest failure mode matters as much as the hit.
Inspiration
At 3am, the fix that would resolve the incident you're staring at usually already exists — in a Slack thread, a stale postmortem, or a teammate's head. Incident knowledge is tribal and ephemeral, so the same failure class gets re-diagnosed from scratch every few weeks.
The obvious move is "give the agent a memory." The interesting question is what that memory has to be. If an agent's memory is a log file or a cache, it degrades exactly when you need it — mid-incident, mid-failover. We wanted memory that is the system of record: transactional, durable, queryable semantically, and unavailable to go down independently of the thing it's remembering for.
What it does
DejaView triages incoming incidents and answers two questions with evidence: have we seen this before, and what actually fixed it?
A new incident is embedded and vector-searched against every past incident and resolution. The agent proposes a diagnosis and remediation that cite the prior incident and its similarity score — never a bare assertion. A live run recalls a 9-day-old precedent at 97% cosine similarity in 2.5ms and proposes the fix that genuinely worked then.
When a human resolves an incident, that resolution is embedded and written straight back into the same table — so the next similar incident can recall it. The loop closes: every resolution makes the next triage smarter.
It also stays honest. When nothing clears the similarity threshold, it says so rather than inventing a match.
How we built it
CockroachDB is the only datastore — and it holds two different kinds of memory:
Semantic memory: a VECTOR(1024) column with a distributed vector index over incidents and resolutions. Recall is ORDER BY embedding <=> $1 LIMIT k. The agent's own execution memory: every step of the Bedrock Converse tool-use loop commits an atomic checkpoint — model I/O, tool call, result, state change — keyed (run_id, step_index). Kill the process mid-triage and the next invocation rebuilds the exact conversation from CockroachDB and continues, guarded by a run lease so at-least-once redelivery can't double-spend on Bedrock.
AWS: Bedrock Titan Embed v2 for embeddings, Claude via Converse for reasoning, Lambda for execution, S3+CloudFront for the dashboard.
A React dashboard makes the recall visible — similarity scores, the recalled resolution, and the agent's reasoning timeline, because the evidence should never be hidden behind a claim.
Challenges we ran into
The vector_cosine_ops trap. CREATE VECTOR INDEX with no opclass silently builds an L2 index. Cosine queries against it still return correct values — they just full-scan. Correctness testing will never catch this; only EXPLAIN will.
The Bedrock us. prefix. anthropic.claude-sonnet-4-6 fails with "on-demand throughput isn't supported." The invocable id is us.anthropic.claude-sonnet-4-6 — a cross-region inference profile. The catalog listing is not the invocable id.
A wedged-run bug class. The first loop only handled lease-lost failures. Anything else — a Bedrock throttle, a malformed response — escaped as a bare exception, leaving the run running with the lease still held. Nothing could reclaim it, and Lambda faithfully replayed the identical deterministic crash three times before giving up silently. The fix needed all of: guarding the whole step body, a terminal failed checkpoint, a DLQ, and an EventBridge sweeper for containers that die with no exception at all.
Fix text doesn't rank on symptom similarity. Our biggest retrieval insight. A resolution ("raised max_connections, added a circuit breaker") is semantically unlike the symptom that finds it, so procedural memories ranked outside top-k — 3 of 5 live incidents wrongly reported "no recorded fix" while the fix sat at rank 9. Raising top_k doesn't help; it just admits unrelated fixes. The answer was to stop hoping and follow the source_incident_id link instead: recall similar episodes, then fetch what was done about them.
Claude access was pending all build. So we ordered the work so every milestone — agent loop, resumability, concurrency, recall — verifies against a deterministic client with an identical tool-calling contract. Titan was unaffected.
Accomplishments that we're proud of
We measured whether memory helps instead of asserting it. make baseline runs the same incident through the same loop twice, with search_memory genuinely withheld from the model's toolConfig in the control arm — not a cooperative fake declining to call it. Memory arm cites the precedent at 0.9683 and quotes the historical fix; the control cites nothing.
Resumability is proven, not claimed. Killed mid-step, resumed by a fresh holder reading only CockroachDB checkpoints, with zero duplicate side-effects and no replayed model calls.
We kept the scoreboard honest. 2 of 4 CockroachDB tools genuinely used, stated plainly rather than padded. The ccloud CLI and MCP server aren't claimed because they weren't used.
What we learned
A green test suite tells you almost nothing about whether the product works. Every significant defect we found came from re-verification or driving a real browser — none from 348 passing tests: a null field that hung the recall panel forever, primary buttons dead in the exact state every run ends in, remediation text that was actually the symptom restated, and a CI secret-scan gate that was red while our own docs claimed it verified green.
Claims outrun implementations, including your own. One review found a CRITICAL whose fix was asserted in the docs but absent from the code. We started treating "verified" as meaning a command was run and its output pasted, and it changed what we caught.
Failure modes are product surface. "No precedent above threshold" needed to be a designed, honest state — not a blank panel or, worse, a low-confidence match dressed up as a hit.
What's next for DejaView Multi-region CockroachDB — the memory layer's actual thesis, currently single-region. Real alerting integrations — PagerDuty/Opsgenie ingestion instead of a chaos simulator. Memory consolidation — dedupe near-identical incidents and decay stale resolutions so recall quality holds as the corpus grows. Gated auto-remediation — today it proposes and a human applies; the checkpointing and audit trail already exist to make a narrow, approval-gated auto-fix path safe. Finish the tool sweep — ccloud provisioning and the MCP server, both prepared but not used.
Built With
- amazon-bedrock
- amazon-web-services
- aws-lambda
- aws-sam
- claude
- cloudformation
- cloudfront
- cockroachdb
- github-actions
- pgvector
- playwright
- psycopg
- pytest
- python
- react
- sql
- typescript
- vector-search
- vite
Log in or sign up for Devpost to join the conversation.