MaaS v2.0 — Institutional Code Memory & On-Call Copilot

Tagline: Automated Rails incident diagnosis powered by zero-token stack trace pre-filtering, Paritok GPU context compression, pgvector institutional memory, and AgentKit Human-in-the-Loop governance.


Inspiration

Every senior Ruby on Rails developer carries an immense amount of institutional memory in their head: Which gem version broke the ActiveRecord charge job last October? Why did PostgreSQL throw a duplicate key violation on stripe_charge_id? Which callback in the user model triggers silent deadlocks?

When production exceptions strike, on-call engineers are inundated with massive, verbose Rails development and production logs (50KB–200KB+ / 100K+ characters per exception trace).

The Massive Rails Developer Problem:

  1. Drowning in Framework Stack Noise: 90% of a Rails stack trace consists of framework gem overhead (sidekiq, puma, active_record, action_pack, rack, puma, concurrent-ruby) rather than application business logic (app/ or lib/).
  2. Exorbitant LLM Token Bills: Dumping raw Rails logs into modern LLMs (Gemini, Claude 3.5, GPT-4o) burns 30,000 to 50,000 input tokens per single incident. For active Rails applications processing dozens of error alerts daily, LLM API bills quickly reach thousands of dollars per month.
  3. Loss of Team Knowledge: When senior engineers fix complex Rails bugs, their rationale is lost in closed PRs or Slack messages, forcing the team to re-diagnose identical incidents months later.

We were inspired to build MaaS v2.0 (On-Call Copilot) to solve both problems at once: turn team postmortems into semantic vector memories while using Paritok GPU Context Compression to make AI-driven Rails incident diagnostics financially viable and lightning fast.


What it does

MaaS v2.0 is an intelligent On-Call Copilot for Ruby on Rails engineering teams. It receives raw production logs or stack traces via web form, drag-and-drop file upload, or JSON API, and executes a multi-layer diagnostic pipeline:

  1. Zero-Token Stack Trace Pre-Filtering: Runs LogNormalizer inside pure Ruby RAM at 0 token cost ($0.00), collapsing repetitive log lines and stripping third-party gem frames (retaining only app/ and lib/ call sites). This immediately reduces log volume by 48.9% to 90.1%.
  2. pgvector Institutional Memory Recall: Performs HNSW cosine similarity vector search over PostgreSQL 16 embeddings to retrieve the top 5 past team postmortems and architectural decisions.
  3. Paritok GPU Context Compression: Sends the prompt bundle to the Paritok Cloud GPU Proxy (https://www.paritok.com/api/compress), squeezing the remaining text semantically by an additional ~70.7%.
  4. Gemini 2.5 LLM Diagnosis: Invokes Gemini 2.5 Flash with the hyper-compressed context to generate a root-cause explanation, cite past team precedents, and provide step-by-step code solutions.
  5. AgentKit HITL Review Inbox: Routes all proposed solutions to a Human-in-the-Loop review console (/agentkit/suggestions) where engineers can Approve, Snooze, or Reject with structured feedback. Rejected suggestions automatically become new institutional memories to prevent repeated mistakes.
  6. AST Class Visualizer & Telemetry: Generates interactive D3 class dependency graphs (/rubrowser) and streams real-time token/dollar savings stats via ActionCable to the dashboard (/agentkit/factory and /architecture).

How we built it

  • Framework & Core: Built on Rails 8.1 and Ruby 3.4+, leveraging AgentKit-Rails 0.2 as the underlying agent orchestration kernel (recall!, build_context, suggest!, and Agentkit::Flow).
  • Context Compression Engine: Integrated Paritok (https://www.paritok.com) via MaaS::ParitokGateway with automatic health checks, real-time ActionCable ledger streaming, and transparent fallback to direct Gemini execution if offline.
  • Zero-Token Pre-Filter: Engineered Oncall::LogNormalizer in pure Ruby using Regex frame matching (^app/ and ^lib/) and consecutive line deduplication.
  • Vector Database: Configured PostgreSQL 16 with pgvector using an HNSW cosine distance index for sub-millisecond retrieval of agentkit_memories.
  • LLM Integration: Connected Gemini 2.5 Flash (gemini-flash-latest) via RubyLLM for structured response generation.
  • AST Dependency Graphing: Built MaaS::RubrowserGenerator using Ruby Parser AST introspection to generate dynamic HTML/D3 dependency visualizations.
  • Architecture & System Views: Designed a custom SVG system architecture page (/architecture) mapping the 6-stage lifecycle, 4-lane topology, context reduction grid, and database schema.

Challenges we ran into

  1. Differentiating App Code from Gem Paths: Naive stack trace regexes often misclassified gem files containing substrings like /lib/ (e.g. .../gems/sidekiq-7.3.1/lib/sidekiq/...) as application code. We solved this by strictly anchoring OWN_FRAME matching to paths starting with app/ or lib/.
  2. Handling Massive Web Payloads: Standard Rack requests can fail or time out on 100KB–200KB log pastes. We added multipart file stream handling (params[:log_file]) and JSON body parsing to support log files of any size smoothly.
  3. Engine vs. Host App Route Resolution: Mounting Agentkit::Engine at /agentkit initially caused routing helper scope mismatches on custom pages like /architecture. We resolved this by building robust URI path helpers that work across both main app and engine controllers.
  4. Ensuring Zero Diagnostic Quality Loss During Compression: Compressing text too aggressively can erase vital exception signatures. By pairing LogNormalizer's deterministic signature preservation with Paritok's semantic GPU compression, we achieved a 97% size reduction while retaining 100% diagnostic accuracy.

Accomplishments that we're proud of

  • 97.1% Total Context Reduction: Successfully reduced a real 123 KB Rails development log (example/development.log) from 122,702 characters (~30,600 tokens) down to ~3,500 characters (~875 tokens).
  • Proven Financial Savings with Paritok: Slashed the cost per incident diagnosis from $0.153 USD down to $0.004 USD—saving $148.63 USD per 1,000 incident calls.
  • Zero Token Waste: Proved that pure Ruby reflection and regex normalization can eliminate 90% of prompt bloat at $0.00 cost before making network requests.
  • Full Operational Transparency: Built a live SVG Architecture page (/architecture), Factory control panel (/agentkit/factory), and Rubrowser AST visualizer (/rubrowser).
  • Complete End-to-End Test Suite & Video Walkthrough: Created an automated 1080p video generation script (maas_oncall_demo.rb) and 1-to-1 synchronized narration scripts (maas_oncall_speech.txt and maas_oncall_speech_es.txt).

What we learned

  • Context is Not Free: Modern LLMs support 1M+ token context windows, but sending uncompressed logs is financially unsustainable. Context compression is the most critical architectural decision for production AI agents.
  • Paritok is an Essential Tool for AI Cost Savings: Integrating Paritok GPU Compression into an LLM pipeline provides an instant, massive ROI without requiring custom model retraining or prompt degradation.
  • Hybrid Multi-Stage Reduction Works Best: Deterministic 0-token code filtering + semantic GPU compression yields far better results than relying on either technique alone.
  • Institutional Memory Closes the Feedback Loop: Combining vector recall with Human-in-the-Loop rejections transforms an AI agent from a static prompter into a self-improving system that learns team conventions over time.

What's next for MaaS v Institutional Code Memory

  • Automated GitHub PR Patch Creation: Enable accepted HITL suggestions to automatically open formatted Pull Requests on GitHub using MaaS::GithubClient.
  • Slack & PagerDuty Webhook Integration: Trigger background incident diagnoses automatically upon receiving PagerDuty or Sentry exception webhooks.
  • Multi-Repo Cross-Project Memory Sync: Expand pgvector memory sharing across microservice repositories so team lessons learned in one Rails service immediately benefit other services.
  • Edge Paritok Proxy Deployment: Package the Paritok proxy into a lightweight sidecar container for local development environments and CI/CD pipelines.

Built With

Share this project:

Updates

Submission history