Inspiration

AI agents are increasingly deployed in autonomous workflows, but they fail quietly—hallucinations, breaking traces, and runtime API errors often go undetected until systems break completely. We wanted to build a safety net that doesn't just log these errors, but acts as a dynamic supervisor. ChitinAI was born out of the need for an enterprise-grade meta-agent that monitors, auto-heals, and optimizes live LLM agents in real time without requiring constant human intervention.

What it does

ChitinAI operates as an autonomous, self-healing meta-agent designed to ensure high reliability and cost efficiency for deployed AI workloads:

  • Real-time Trace Scanner (Watcher): Monitors execution flows of live LLM agents, capturing step-level traces and detecting hallucinations or logic breaches.
  • Vector Memory & Auto-Healing: Leverages CockroachDB Distributed Vector Search to fetch historical incident data and triggers ccloud CLI scripts to heal broken application states automatically.
  • Smart Cost Routing: Dynamically routes model requests between high-capability models (Claude 3.5 Sonnet) and lightweight alternatives (Amazon Nova) via AWS Bedrock to drastically reduce API operational costs.
  • Visual Cockpit: Features a custom React/Vite dashboard providing real-time prompt diffs, execution trajectory replays, and postmortem alerts to Slack/Discord.

How we built it

We architected ChitinAI using a robust modular stack tailored for scale and execution speed:

  • Backend & Agent Logic: Core pipeline written in Python, featuring custom watcher threads and orchestration logic.
  • Vector RAG & Memory: CockroachDB with Bedrock Titan embeddings for fast similarity searches across failure postmortems.
  • Infrastructure & Routing: AWS Bedrock SDK for model switching and automated database CLI commands (ccloud) for execution context repairs.
  • Frontend Cockpit: React with Vite for low-latency visual log monitoring and prompt diff displays.
  • Devin Integration: Utilized Devin to audit our core codebase, refactor edge-case error handlers in chitin/watcher.py, and auto-generate comprehensive unit test suites using pytest.

Challenges we ran into

  • Sub-second Fault Detection: Ingesting live agent traces and determining failures without adding latency to the main execution loop required optimizing our background thread orchestration.
  • State Reconstruction during Auto-Healing: Mapping vector search results to actionable CLI recovery steps demanded careful error-type categorization and fail-safe fallbacks.
  • Dynamic Model Switching: Balancing token usage and latency thresholds while seamlessly hot-swapping between Claude 3.5 Sonnet and Amazon Nova inside AWS Bedrock without dropping prompt context.

Accomplishments that we're proud of

  • Zero-Downtime Recovery Loop: Successfully demonstrated end-to-end self-healing where a failing agent context was detected, matched against vector memory, and auto-patched in real time.
  • Significant Cost Reduction: Reduced token spend during benchmark evaluation runs by using dynamic model routing via AWS Bedrock.
  • End-to-End Production Readiness: Built a complete full-stack product featuring backend pipelines, live vector database integration, deployment on Replit, and an intuitive UI dashboard.

What we learned

  • Multi-agent architectures require strict isolation between the "worker" and the "supervisor" to avoid recursive failure loops.
  • Combining vector memory RAG with command-line orchestration unlocks a powerful framework for infrastructure-level self-healing.
  • Agentic workflows built alongside autonomous AI developers like Devin can dramatically accelerate testing, bug discovery, and code refactoring cycles.

What's next for ChitinAI

  • Multi-Cloud Healing Drivers: Expand auto-healing connectors to support Kubernetes, Vercel, and AWS Lambda deployments.
  • Open-Telemetry (OTel) Native Integration: Support standardized tracing protocols so developers can connect ChitinAI to existing monitoring stacks (Datadog, LangSmith, Honeycomb) out of the box.
  • Proactive Anomaly Prediction: Train predictive models on failure traces to patch prompt drifts before a failure even occurs.

Built With

Share this project:

Updates

Submission history