-
-
EdgeOps Agent: Cognitive SRE with Distributed Vector Memory
-
System Architecture: Bridging physical hardware (ESP32) and cognitive LLM-driven SRE via MQTTS and CockroachDB vector memory.
-
Safety First: Real-time Slack Socket Mode integration for human approval before executing destructive actions.
-
Live Execution: The EdgeOps Agent autonomously ingesting live hardware telemetry and deducing the root cause.
-
The Cognitive Brain: LangGraph state machine with an interactive Human-In-The-Loop (HITL) execution gate.
-
Distributed Vector Memory: Using pgvector on CockroachDB Serverless to store and retrieve historical incidents with temporal decay.
-
Fleet Observability: Real-time Prometheus and Grafana dashboards tracking autonomous anomaly resolutions across the edge.
-
AI Traceability: LangSmith observability capturing deep execution traces and token usage across the tiered LLM reasoning pipeline.
-
Seeed XIAO ESP32-S3 edge node with I2C OLED displaying live MQTTS connection status and telemetry ping countdown
-
GIF
Physical SRE in action: ESP32 node reporting "Thermal Runaway" over MQTTS & executing the agent's live "FAN_ON" remediation command.
EdgeOps Agent: Cognitive SRE with Distributed Vector Memory
📺 Demo Video: https://youtu.be/JpVPppSWOFw
💻 GitHub Repository: https://github.com/shawnsony07/Autonomous-SRE
Inspiration
While software-based Site Reliability Engineering (SRE) agents are becoming common for restarting cloud pods or rolling back deployments, we asked ourselves: what happens when physical infrastructure fails?
There is a massive philosophical gap between deterministic control (dumb local if/else rules) and true cognitive SRE. A local microcontroller can easily execute if (temp > 85) { turn_on_fan(); }. But this only treats the symptom, not the root cause. If a bad OTA update introduced an infinite loop pinning the CPU and generating heat, turning on a fan wastes battery and masks the real issue.
Furthermore, isolated edge nodes lack global context. If 50 nodes in a server rack turn on cooling fans simultaneously, the massive current draw could brown out the entire swarm.
We were inspired to build EdgeOps Agent, a cognitive engine for when local if/else statements fail. It bridges the gap between physical edge environments and LLM-driven AI by analyzing global telemetry, learning from past physical incidents via vector memory, and healing the fleet autonomously.
What it does
EdgeOps Agent is a fully autonomous, LLM-powered SRE designed for the physical edge.
- Detects & Ingests: It monitors hardware telemetry (e.g., from an ESP32) via secure MQTTS.
- Retrieves Memory: It queries CockroachDB using an Approximate Nearest Neighbor (ANN) vector search to find similar historical incidents and their successful resolutions.
- Reasons & Plans: Using a tiered LLM approach (Gemini 3.6/3.5 Flash and Gemini 3.5/3.1 Flash Lite via LiteLLM, with Llama 3.2 1B as a local fallback), it deduces the root cause (e.g., software lockup vs. ambient heat) and formulates a remediation plan.
- Executes & Heals: It executes Model Context Protocol (MCP) tools to query databases or sends MQTT commands back to the edge (e.g.,
THROTTLE_CPU,RESET_I2C) to fix the issue. - Human-In-The-Loop (HITL): For critical or destructive actions (like dropping a corrupted table), it pauses execution and requests human permission via an interactive Slack Socket Mode integration.
Live end-to-end agent execution capturing live telemetry and resolving an anomaly.
Slack Socket Mode HITL notification for destructive actions.
How we built it
The architecture is a highly asynchronous, event-driven pipeline bridging the physical and digital worlds:

- The Edge Node: A physical Seeed Studio XIAO ESP32-S3 microcontroller written in C++ (Arduino). It publishes telemetry and receives commands over an encrypted MQTTS broker (port 8883). We integrated an OLED display so you can physically see the hardware state changes and watch how the agent communicates with the edge in real time. (If you don't have the hardware, you can use the included
scripts/mock_alerts.pyscript to simulate anomalies).

Physical hardware setup with the Seeed Studio XIAO ESP32-S3 and 0.96" OLED display.

OLED animation showing the hardware reacting to a FAN_ON remediation command sent by the SRE agent.

The physical ESP32 executing remediation commands locally (Arduino IDE Serial Monitor).
- The Brain: A LangGraph state machine written in Python, compiled with
interrupt_beforehooks to enable the HITL gate.

Cloud Infrastructure (AWS): Hosted the Dockerized agent, MQTT broker, and observability suite on an AWS EC2 instance, injected credentials dynamically using AWS Secrets Manager, and continuously archived pruned vector memories to Amazon S3.
Vector Memory: We used CockroachDB Serverless with the
pgvectorextension to serve as the agent's long-term memory. It stores 768-dimensional embeddings of all past incidents.
CockroachDB executing distributed vector similarity search.
Live vector embeddings stored securely in CockroachDB.
- LLM Routing: A self-hosted LiteLLM gateway securely routes requests to Google Gemini, implementing rate limiting, tiered reasoning, and circuit breaking.
- Tool Execution: We integrated the Model Context Protocol (MCP) directly with CockroachDB to allow the LLM to inspect database schemas and run queries natively.
Challenges we ran into
- Vector Memory Decay: Raw vector similarity searches often surfaced outdated resolutions that no longer applied. We solved this by writing a custom CTE SQL query in CockroachDB that applies a chronological temporal decay penalty to older embeddings, ensuring the agent prioritizes recent learnings.
- MCP Transport Roadblocks: The managed CockroachDB MCP server strictly required HTTP POST requests, but the standard
sse_clientwe started with used GET. We had to dig deep into the MCP spec and implement a customstreamable_httpclient to ensure seamless LLM-to-Database tool execution. - Async Exception Unmasking: Using Python 3.11's
asyncio.TaskGroupwas incredible for performance, but it wrapped nested tool errors inExceptionGroupobjects. This masked the true root causes from the LLM. We had to write a custom traceback unwrapper to expose the inner MCP errors so the agent could learn and retry effectively. - State Machine Halting for HITL: We needed to pause the python event loop for human approval without stalling the MQTT ingestion queue. We built a dual-channel race condition using
asyncio.wait(FIRST_COMPLETED) to race terminal input against real-time WebSocket events from Slack.
Accomplishments that we're proud of
- True Cognitive SRE: Building a system that actually transcends simple automation. Because the agent leverages CockroachDB vector memory, it can "remember" that the last three times it turned on a fan during a specific anomaly, the I2C bus crashed—and autonomously choose a different remediation path.
- Flawless Observability: We integrated Prometheus, Grafana, and LangSmith. Being able to watch the LangGraph waterfall trace, monitor P99 latency spikes during deep reasoning, and track exact token costs per incident gave us total visibility into the agent's brain.
Deep execution tracing of the LangGraph state machine via LangSmith.
Token usage tracking across reasoning tiers.
Prometheus & Grafana live observability tracking anomaly resolutions.
- Safety First: Implementing a Dead Letter Queue (DLQ) in CockroachDB for failed remediations, and a cron job that safely archives stale vector memories to AWS S3 before pruning the database, ensuring long-term performance and data safety.
What we learned
- Swarm Behavior Matters: Designing for the physical edge requires acknowledging swarm behavior. The realization that thousands of isolated edge nodes making logical "local" decisions can actually crash the entire global network was a massive paradigm shift.
- Memory Needs Maintenance: Vector databases are incredible, but they require hygiene. Pruning stale embeddings and applying temporal decay drastically improved the LLM's accuracy and hallucination rates.
- The Power of the MCP Ecosystem: The Model Context Protocol is the future of agentic tooling. Giving the LLM a standardized, typed interface to CockroachDB completely eliminated the need for us to write brittle custom database proxy scripts.
What's next for EdgeOps Agent: Cognitive SRE with Distributed Vector Memory
While the current architecture proves the viability of Cognitive SRE, there are several limitations we need to address to make the system truly robust for enterprise-grade production:
- Handling Network Partitions (Offline Resilience): The agent currently assumes a reliable MQTT connection. If an edge node loses network access, the central agent is blind. We need to implement a localized fallback state machine on the hardware itself so the ESP32 can safely degrade when disconnected from the global brain.
- Multi-Agent Orchestration: A single SRE agent will bottleneck when managing thousands of nodes. We plan to scale to a multi-agent setup where specialized agents (e.g., Database SRE, Network SRE, Edge SRE) handle their respective domains and report up to a master orchestrator.
- Autonomous Verification (
Verify_Fix): Right now, the agent assumes its executed action succeeded. We need to implement a verification loop in the LangGraph to autonomously confirm that the hardware telemetry actually returned to normal after a remediation was applied. - Expanded MCP Integrations: Plugging the agent directly into AWS EC2 and VPC networking MCP servers, allowing it to cycle power on degraded bare-metal servers or failover misbehaving switch ports natively without intermediate proxies.
Built With
- amazon-web-services
- arduino
- aws-ec2
- cockroachdb
- docker
- esp32
- google-gemini
- grafana
- langchain
- langgraph
- langsmith
- litellm
- mcp
- mqtt
- ollama
- pgvector
- prometheus
- python
- slack


Log in or sign up for Devpost to join the conversation.