MaaS v2.0 — Institutional Code Memory & On-Call Copilot
Tagline: Automated Rails incident diagnosis powered by zero-token stack trace pre-filtering, Paritok GPU context compression, pgvector institutional memory, and AgentKit Human-in-the-Loop governance.
Inspiration
Every senior Ruby on Rails developer carries an immense amount of institutional memory in their head: Which gem version broke the ActiveRecord charge job last October? Why did PostgreSQL throw a duplicate key violation on stripe_charge_id? Which callback in the user model triggers silent deadlocks?
When production exceptions strike, on-call engineers are inundated with massive, verbose Rails development and production logs (50KB–200KB+ / 100K+ characters per exception trace).
The Massive Rails Developer Problem:
- Drowning in Framework Stack Noise: 90% of a Rails stack trace consists of framework gem overhead (
sidekiq,puma,active_record,action_pack,rack,puma,concurrent-ruby) rather than application business logic (app/orlib/). - Exorbitant LLM Token Bills: Dumping raw Rails logs into modern LLMs (Gemini, Claude 3.5, GPT-4o) burns 30,000 to 50,000 input tokens per single incident. For active Rails applications processing dozens of error alerts daily, LLM API bills quickly reach thousands of dollars per month.
- Loss of Team Knowledge: When senior engineers fix complex Rails bugs, their rationale is lost in closed PRs or Slack messages, forcing the team to re-diagnose identical incidents months later.
We were inspired to build MaaS v2.0 (On-Call Copilot) to solve both problems at once: turn team postmortems into semantic vector memories while using Paritok GPU Context Compression to make AI-driven Rails incident diagnostics financially viable and lightning fast.
What it does
MaaS v2.0 is an intelligent On-Call Copilot for Ruby on Rails engineering teams. It receives raw production logs or stack traces via web form, drag-and-drop file upload, or JSON API, and executes a multi-layer diagnostic pipeline:
- Zero-Token Stack Trace Pre-Filtering: Runs
LogNormalizerinside pure Ruby RAM at 0 token cost ($0.00), collapsing repetitive log lines and stripping third-party gem frames (retaining onlyapp/andlib/call sites). This immediately reduces log volume by 48.9% to 90.1%. - pgvector Institutional Memory Recall: Performs HNSW cosine similarity vector search over PostgreSQL 16 embeddings to retrieve the top 5 past team postmortems and architectural decisions.
- Paritok GPU Context Compression: Sends the prompt bundle to the Paritok Cloud GPU Proxy (
https://www.paritok.com/api/compress), squeezing the remaining text semantically by an additional ~70.7%. - Gemini 2.5 LLM Diagnosis: Invokes Gemini 2.5 Flash with the hyper-compressed context to generate a root-cause explanation, cite past team precedents, and provide step-by-step code solutions.
- AgentKit HITL Review Inbox: Routes all proposed solutions to a Human-in-the-Loop review console (
/agentkit/suggestions) where engineers can Approve, Snooze, or Reject with structured feedback. Rejected suggestions automatically become new institutional memories to prevent repeated mistakes. - AST Class Visualizer & Telemetry: Generates interactive D3 class dependency graphs (
/rubrowser) and streams real-time token/dollar savings stats via ActionCable to the dashboard (/agentkit/factoryand/architecture).
How we built it
- Framework & Core: Built on Rails 8.1 and Ruby 3.4+, leveraging AgentKit-Rails 0.2 as the underlying agent orchestration kernel (
recall!,build_context,suggest!, andAgentkit::Flow). - Context Compression Engine: Integrated Paritok (
https://www.paritok.com) viaMaaS::ParitokGatewaywith automatic health checks, real-time ActionCable ledger streaming, and transparent fallback to direct Gemini execution if offline. - Zero-Token Pre-Filter: Engineered
Oncall::LogNormalizerin pure Ruby using Regex frame matching (^app/and^lib/) and consecutive line deduplication. - Vector Database: Configured PostgreSQL 16 with pgvector using an HNSW cosine distance index for sub-millisecond retrieval of
agentkit_memories. - LLM Integration: Connected Gemini 2.5 Flash (
gemini-flash-latest) via RubyLLM for structured response generation. - AST Dependency Graphing: Built
MaaS::RubrowserGeneratorusing Ruby Parser AST introspection to generate dynamic HTML/D3 dependency visualizations. - Architecture & System Views: Designed a custom SVG system architecture page (
/architecture) mapping the 6-stage lifecycle, 4-lane topology, context reduction grid, and database schema.
Challenges we ran into
- Differentiating App Code from Gem Paths: Naive stack trace regexes often misclassified gem files containing substrings like
/lib/(e.g..../gems/sidekiq-7.3.1/lib/sidekiq/...) as application code. We solved this by strictly anchoringOWN_FRAMEmatching to paths starting withapp/orlib/. - Handling Massive Web Payloads: Standard Rack requests can fail or time out on 100KB–200KB log pastes. We added multipart file stream handling (
params[:log_file]) and JSON body parsing to support log files of any size smoothly. - Engine vs. Host App Route Resolution: Mounting
Agentkit::Engineat/agentkitinitially caused routing helper scope mismatches on custom pages like/architecture. We resolved this by building robust URI path helpers that work across both main app and engine controllers. - Ensuring Zero Diagnostic Quality Loss During Compression: Compressing text too aggressively can erase vital exception signatures. By pairing
LogNormalizer's deterministic signature preservation with Paritok's semantic GPU compression, we achieved a 97% size reduction while retaining 100% diagnostic accuracy.
Accomplishments that we're proud of
- 97.1% Total Context Reduction: Successfully reduced a real 123 KB Rails development log (
example/development.log) from 122,702 characters (~30,600 tokens) down to ~3,500 characters (~875 tokens). - Proven Financial Savings with Paritok: Slashed the cost per incident diagnosis from $0.153 USD down to $0.004 USD—saving $148.63 USD per 1,000 incident calls.
- Zero Token Waste: Proved that pure Ruby reflection and regex normalization can eliminate 90% of prompt bloat at $0.00 cost before making network requests.
- Full Operational Transparency: Built a live SVG Architecture page (
/architecture), Factory control panel (/agentkit/factory), and Rubrowser AST visualizer (/rubrowser). - Complete End-to-End Test Suite & Video Walkthrough: Created an automated 1080p video generation script (
maas_oncall_demo.rb) and 1-to-1 synchronized narration scripts (maas_oncall_speech.txtandmaas_oncall_speech_es.txt).
What we learned
- Context is Not Free: Modern LLMs support 1M+ token context windows, but sending uncompressed logs is financially unsustainable. Context compression is the most critical architectural decision for production AI agents.
- Paritok is an Essential Tool for AI Cost Savings: Integrating Paritok GPU Compression into an LLM pipeline provides an instant, massive ROI without requiring custom model retraining or prompt degradation.
- Hybrid Multi-Stage Reduction Works Best: Deterministic 0-token code filtering + semantic GPU compression yields far better results than relying on either technique alone.
- Institutional Memory Closes the Feedback Loop: Combining vector recall with Human-in-the-Loop rejections transforms an AI agent from a static prompter into a self-improving system that learns team conventions over time.
What's next for MaaS v Institutional Code Memory
- Automated GitHub PR Patch Creation: Enable accepted HITL suggestions to automatically open formatted Pull Requests on GitHub using
MaaS::GithubClient. - Slack & PagerDuty Webhook Integration: Trigger background incident diagnoses automatically upon receiving PagerDuty or Sentry exception webhooks.
- Multi-Repo Cross-Project Memory Sync: Expand pgvector memory sharing across microservice repositories so team lessons learned in one Rails service immediately benefit other services.
- Edge Paritok Proxy Deployment: Package the Paritok proxy into a lightweight sidecar container for local development environments and CI/CD pipelines.
Built With
- on
- paritok
- postgresql
- ruby
- ruby-on-rails
Log in or sign up for Devpost to join the conversation.