Inspiration

Every on-call engineer has lived the same 3 a.m. moment: an alert fires, and before you can fix anything you have to remember. Have we seen this before? What actually caused it last time? What changed recently? That knowledge exists — but it's scattered across old incident write-ups, a wiki nobody reads, merge history, dashboards, and the code itself. So the first 40 minutes of every incident is spent rediscovering what the team already learned.

We wanted to build an on-call teammate whose entire value is memory — and to prove a specific claim: for incident response, memory makes an agent useful. A frontier LLM with no context is just a confident stranger. Give it the team's accumulated experience, indexed by meaning and wired to the code, and it becomes the senior engineer who's seen this before.

What it does

felix takes an incoming alert and, before saying anything, recalls by meaning everything the organization already knows that's relevant, then reasons over it to a diagnosis with a concrete fix, and finally writes the resolution back so the next incident is easier.

Memory Recalled by
Past incidents & resolutions Vector search
Project documentation Vector search
Recent merges Vector search + time window ("what changed lately?")
Curated runbooks Vector search by trigger text
Code graph (call relationships) Recursive-CTE graph traversal
Service topology & live metrics Recursive walk → live metrics over downstream services

Around that core, it's a five-panel web app:

  • Triage: Alert → live-streamed diagnosis with an evidence panel linking every citation back to its source.
  • Incident Library: Semantic search over historical incidents.
  • Live Monitoring: Real-time stream fed by a CockroachDB changefeed.
  • DB Overview: Powered end-to-end by the CockroachDB Cloud Managed MCP Server.
  • ccloud Terminal: An interactive, real in-browser terminal experience.
  • Onboard: Onboard a new service for felix to act as an SRE assistant for.

How we built it

CockroachDB is the entire memory layer, heavily utilizing:

  • Native VECTOR(1024) type & vector indexes for semantic recall. Every incident, doc, merge, and runbook is embedded at load time.
  • Recursive CTEs (WITH RECURSIVE) over an edge table for the code graph. The parser turns the sample service into code_nodes and code_edges (directed in the call direction).
  • CDC Changefeed on the metrics table driving the live monitoring panel and a watcher service that raises an alert the instant an anomaly appears.

Challenges we ran into

  • Parsing code into queryable graph edges: Building a code parser that reliably maps callers and callees into code_nodes and code_edges while handling service routes.
  • Closed-loop CDC alert streaming: Connecting the live sample service's metric stream to CockroachDB changefeeds, maintaining open server connections without connection leaks, and streaming continuous data to felix for real-time alert generation.
  • Evidence scoring & multi-source reranking: Digging through fragmented knowledge sources (past incidents, runbooks, code graph paths, project docs, and recent merges) and scoring candidate records so Gemini can present users with a transparent, ranked evidence panel showing why each piece was deemed relevant.

Accomplishments that we're proud of

  • Built a complete, memory-first SRE agent where CockroachDB handles vector search, graph traversal, and CDC events in a single database.
  • Designed an explainable evidence panel where engineers can inspect every past incident, code path, and doc the LLM cited, complete with relevance scores that ground the LLM's diagnosis in verifiably real history.
  • Created a deterministic "symptom-to-cause" tracing pipeline using recursive CTEs that walks upstream through code dependencies and topology graphs to identify the root cause component before invoking the LLM.

What we learned

We learned how much CockroachDB can carry on its own. Vector search, recursive graph traversal, and change data capture are usually three separate systems; having them in one database — one connection string, one place where memory lives and compounds — made the memory-based architecture genuinely simple instead of a complex distributed-systems project.

What's next for felix

Today felix ships with one carefully-built demo service. The next step is making it plug-and-play for any production environment:

  • Automated Codebase Onboarding: Generalize the parser to automatically ingest any Git repository into the code graph, with connectors for PagerDuty, Confluence, Notion, and GitHub.
  • Bring Your Own Telemetry: Replace sample metric generators with native adapters for Prometheus, Datadog, and OpenTelemetry streams.
  • Multi-Tenant Service Registry: Support multi-service deployments where felix automatically routes alerts to the correct service's isolated code graph and incident history.
  • Closed-Loop Action Execution: Expand beyond recommendations by allowing felix to generate PR fixes and trigger automated rollback pipelines subject to human-in-the-loop approval.

Built With

Share this project:

Updates

Submission history