Inspiration

Production incidents are rarely isolated events. Engineering teams often encounter similar failures repeatedly, yet the knowledge from previous incidents is fragmented across tickets, chat threads, runbooks, postmortems, and individual experience.

We built OpsRelay to address this gap by giving incident-response agents persistent operational memory.

Instead of treating every incident as a new problem, OpsRelay retains relevant historical context, retrieves similar past incidents, and surfaces verified operational evidence to help engineers respond faster and make better-informed decisions.

Our objective was to make memory a core part of the agent architecture, rather than treating the database as passive application storage.

What it does

OpsRelay is an AI-powered incident response and shift-handoff platform designed around persistent agent memory.

Engineers can provide raw incident notes, alerts, or logs. Amazon Bedrock processes this information and extracts structured details such as:

  • Severity
  • Affected service
  • Incident timeline
  • Root cause
  • Recommended actions
  • Follow-up tasks

Before the extracted information becomes part of the organization's trusted knowledge base, it is reviewed and approved by an engineer.

Once approved, the incident becomes part of OpsRelay's long-term memory in CockroachDB.

The platform can then:

  • Retrieve semantically similar historical incidents
  • Recall previous remediation steps and operational context
  • Query verified evidence through CockroachDB Cloud Managed MCP
  • Generate grounded recommendations based on historical data
  • Preserve conversation history and task state across sessions
  • Produce structured incident summaries and shift handoffs

This allows the system to become more valuable as additional incidents and verified operational knowledge are accumulated.

How we built it

OpsRelay uses CockroachDB as the persistent memory layer for the agent system.

Rather than separating transactional data, vector embeddings, conversation state, and verified evidence across multiple storage systems, OpsRelay keeps these components within a distributed PostgreSQL-compatible platform.

CockroachDB Distributed Vector Indexing

After an incident is reviewed and approved, relevant incident information is converted into an embedding using Amazon Titan Embed Text v2.

The resulting 1024-dimensional vectors are stored directly in CockroachDB alongside incident metadata.

When an engineer investigates a new issue, OpsRelay performs semantic vector search across historical incidents to identify similar failures based on symptoms, affected services, and operational patterns.

The retrieved incidents are then supplied to the reasoning agent as contextual memory.

This enables the agent to answer questions such as:

  • Have we encountered this issue before?
  • What resolved a similar incident?
  • Which services were affected in previous occurrences?
  • What investigation steps were effective?
  • Is this part of a recurring failure pattern?

By grounding responses in historical incident data, OpsRelay reduces the agent's dependence on general model knowledge and makes its recommendations more relevant to the organization's actual operating environment.

CockroachDB Cloud Managed MCP Server

OpsRelay also integrates the CockroachDB Cloud Managed MCP Server for an evidence-based investigation workflow.

A key design principle was that AI-generated conclusions should not automatically become trusted organizational knowledge.

The AI can extract potential facts, causes, and recommendations from incident information, but an engineer must validate and approve that information before it is promoted to trusted memory.

Approved facts are stored as verified operational evidence in CockroachDB.

The investigator agent can then query this evidence through Managed MCP in a controlled, read-only workflow and use it to generate grounded responses supported by trusted data.

This creates two complementary forms of memory:

  1. Semantic memory for identifying historically similar incidents through vector search
  2. Verified operational memory for retrieving trusted facts through Managed MCP

Together, these mechanisms allow OpsRelay to retrieve broad historical context while maintaining a clear distinction between generated information and verified organizational knowledge.

Amazon Bedrock

OpsRelay uses multiple Amazon Bedrock models for specialized responsibilities:

  • Claude Haiku 4.5 for extracting structured incident information from raw notes
  • Amazon Nova 2 Lite for reasoning and grounded response generation
  • Amazon Titan Embed Text v2 for vector embedding generation

The application is deployed on AWS EC2 with a React and TypeScript frontend and an Express backend.

CockroachDB stores incidents, tasks, conversations, user relationships, vector embeddings, and verified operational evidence.

Challenges we ran into

One of the most important challenges was determining what information the agent should be allowed to retain as trusted memory.

Automatically storing every model-generated conclusion could introduce incorrect assumptions into future incident investigations. Over time, those inaccuracies could influence subsequent recommendations.

To address this, we introduced a human-in-the-loop approval process.

The AI can extract information and propose conclusions, but only engineer-approved evidence becomes part of the trusted long-term knowledge base.

Another challenge was managing several types of operational state within the same system, including:

  • Transactional incident data
  • Semantic vector memory
  • Conversation history
  • Task state
  • User authorization
  • Verified evidence accessible through MCP

CockroachDB allowed us to support these requirements within a single distributed data platform instead of introducing separate databases for each type of state.

We also had to ensure that memory retrieval respected authorization boundaries. The agent must only retrieve incidents and evidence that the current user is permitted to access.

Accomplishments that we're proud of

One of our primary accomplishments is that CockroachDB directly influences the behavior of the agent rather than serving only as application storage.

OpsRelay uses CockroachDB to:

  • Persist incident history
  • Retrieve semantically related failures
  • Preserve conversation history
  • Maintain task and incident state
  • Store verified operational knowledge
  • Expose trusted evidence to agents through Managed MCP

We are also proud of the platform's human-reviewed memory architecture.

AI-generated information must be validated before it becomes trusted long-term memory. This allows the system to benefit from AI-assisted extraction while maintaining accountability and data quality.

As a result, OpsRelay's operational memory remains both useful and auditable.

What we learned

Building OpsRelay reinforced that the effectiveness of an agentic system depends on much more than the underlying foundation model.

A reliable agent also needs mechanisms to:

  • Retain relevant context
  • Retrieve the correct historical information
  • Distinguish generated content from verified knowledge
  • Respect authorization boundaries
  • Maintain state across sessions

We also learned that vector search and MCP address different but complementary retrieval requirements.

Vector search is effective for discovering relevant historical incidents when similarities are not explicitly defined.

Managed MCP provides a controlled interface for querying structured and trusted operational evidence.

Using both approaches allowed us to build a system that can discover relevant historical context while grounding important conclusions in verified data.

What's next for OpsRelay

Our next goal is to evolve OpsRelay into a more proactive incident-response and operational intelligence platform while keeping engineers in control of critical decisions.

Planned improvements include:

  • Automated incident correlation across services
  • Intelligent engineer routing based on historical expertise
  • Detection of recurring incident patterns
  • Cross-region operational memory
  • Automated runbook recommendations
  • AI-generated postmortems grounded in verified evidence
  • Incident risk scoring using historical failure patterns
  • Deeper CockroachDB observability and ccloud automation

Our long-term vision is for OpsRelay to serve as a persistent operational memory layer for engineering teams, ensuring that knowledge gained from every incident remains available to improve the response to the next one.

Built With

Share this project:

Updates

Submission history