Inspiration
Long horizon systems like Claude Code or Codex depend on context quality. When docs have broken links, stale content, or missing context, these agents hallucinate and waste developer time. Yet while production code is continuously tested and monitored, context knowledge is treated as a static artifact, written once, rarely maintained, and never checked for "AI-readiness." Inspiring from Sentry for application errors, and Prometheus for infrastructure metrics, this project treats knowledge bases as engineering artifacts that should be continuously evaluated and maintained, just like production code.
What it does
The system runs a feedback loop which is diagnose, propose, execute, verify, and learn. Ten deterministic signal collectors across six dimensions (retrieval, context, consistency, trust, connectivity, workflow) scan artifacts and produce signals — broken links, stale content, orphaned pages, missing metadata, poor heading quality, terminology inconsistencies. Problems are then rank (encoding score, outcome history, retrieval score, staleness factor) and generates root cause analysis. LLM then generates remediation proposals with reasoning and rollback plan, guaranteeing signal resolution without introducing new problems.
An auto-loop applies approved modifications, fixing broken links, adding metadata, creating cross-references, updating content, hen the system re-assesses and compares scores. If regression is detected, changes are rolled back automatically. Outcomes (strategy, score change, tokens consumed, full decision traces) are persisted to CockroachDB, and before each new proposal, the agent retrieves its past outcomes to avoid failed strategies and reinforce successful ones.
How I built it
The system is an event-driven knowledge-base maintenance agent built with AWS Lambda, S3, CockroachDB, Apache Burr, and Groq.
Knowledge artifacts are stored as Markdown files in Amazon S3. When a file is uploaded or updated, an API call triggers AWS Lambda, which runs the assessment pipeline. The system analyzes the documentation for problems such as broken links, dangling references, mixed topics, and missing context, then stores the detected signals and artifact relationships in CockroachDB. Once finish, the remediation workflow is orchestrated by Apache Burr and persisted in CockroachDB. Burr tracks the workflow state across Lambda invocations, so a workflow can pause for human approval and resume hours later without losing context.
An LLM then generates remediation proposals. Before accepting the proposal, the system compares each proposal against a deterministic baseline using four constraints: coverage, edit budget, safety, and confidence. After approval, Lambda downloads the affected files into its temporary /tmp storage, applies the selected remediation, and uploads only changed files back to S3. Before modification, the system maintains an in-memory backup, making a failed or harmful remediation can be reverted even if the Lambda container that performed the change no longer exists.
Once the flow is complete, CockroachDB acts as the agent's persistent memory to store the following:
- Working memory: Burr workflow state, allowing workflows to resume across Lambda invocations.
- Long-term memory: remediation history, including strategies, score changes, token usage, and decision traces.
- Semantic memory: 384-dimensional embeddings for knowledge artifacts, enabling similarity search for related documents and previous problems. which set the context for the LLM to retrieve relevant remediation history, allowi it to reuse strategies that worked previously instead of repeatedly solving the same problem from scratch.
I also exposed CockroachDB through its Managed MCP server, providing external AI agents with read-only access to the maintenance agent's accumulated knowledge.
What I learned
Artifact selection dramatically affects results. Simple docs (30 artifacts) behaved very differently from deeper docs (100 artifacts) which makes demo results must be validated across multiple subsets
Deterministic proposal is essential. LLM proposals can worsen results when they mismatch detection formats, exceed edit budgets, or introduce new problems. The LLM should therefore be an enhancement layer that can be rejected, not a replacement for deterministic safeguards.
ACID transactions are foundational for serverless agents. Concurrent Lambda invocations can corrupt shared state without serializable isolation and conflict retries. CockroachDB’s distributed ACID guarantees make the serverless agent architecture reliable and should be designed in from the start.
What's next for AI Knowledge Maintenance
Multi-agent orchestration could leverage the MCP server's agent-to-agent communication. Multiple specialized agents (one for broken links, one for stale content, one for structural issues) could run concurrently, sharing institutional memory via CockroachDB MCP views.
Production deployment with CI/CD integration will be implemented for developer pushes docs to S3, the system automatically assesses and remediates, results are visible via a dashboard, and external agents query the maintenance agent's memory via MCP to inform their own documentation-related decisions.
Built With
- amazon-web-services
- apache-burr
- cockroachdb
- python
Log in or sign up for Devpost to join the conversation.