DriftGate: Before You Trust an Eval, We Test the Test.
Inspiration: The Silent $100k Prompt Bug
Every team shipping LLM applications has lived through this nightmare:
You iterate on a system prompt for a customer support agent. You make it slightly "friendlier" and "more empathetic." You run your standard eval suite: 100% of outputs are valid JSON, and your aggregate benchmark score didn't budge.
You deploy to production.
Three days later, finance messages you in a panic: Account closure tickets are failing silently. Why? In Prompt v2, the model decided that "friendlier" meant renaming the root JSON schema key from "summary" to "description" on account closures only.
Your aggregate evaluations saw 100% valid JSON. Your generic LLM-as-a-Judge gave every response a thumbs up. The test was broken, but you had no way of knowing.
We realized a fundamental truth that the entire AI industry is ignoring: Everyone asks you to trust hand-written evals and LLM judges, but nobody tests the tests.
We built DriftGate to solve this: a self-calibrating regression testing infrastructure that ingests raw production telemetry, automatically partitions behavior into vector clusters, synthesizes deterministic invariants, and audits every verifier against ground-truth golden slices before it is ever allowed to grade your model.
What DriftGate Does
DriftGate turns unstructured production traces into calibrated mathematical test suites and proves whether your new prompt or model caused a behavioral regression.
$$\text{TRACE} \longrightarrow \text{CLUSTER} \longrightarrow \text{SYNTHESIZE} \longrightarrow \text{CALIBRATE} \longrightarrow \text{RUN} \longrightarrow \text{MEASURE}$$
- Autonomous Behavioral Partitioning: Instead of forcing engineers to hand-craft test cases, DriftGate embeds raw multi-turn traces into 1536-dimensional vector space on Neon Serverless Postgres (
pgvector) and discovers natural operational boundaries using HDBSCAN density clustering. - LangGraph Invariant Synthesis: A 4-node LangGraph state machine analyzes historical cluster traces to synthesize deterministic AST assertions, JSON schema constraints, tool trajectory validators, and semantic checks.
- The Calibration Gate (Our Moat): A test that passes broken outputs is worse than no test. DriftGate tests each verifier against golden ground-truth slices. Any rubber-stamp judge with a False Pass Rate ( \text{FPR} > 0.100 ) or Recall ( \text{Recall} < 0.800 ) is hard-blocked and purged from the suite.
- Featherless Multi-Model Runner: Executes parallel sample evaluations across 32,000+ open-source models with SHA-256 result caching.
- Frequentist Statistical Rigor: Computes Wilson 95% Confidence Intervals, two-proportion z-tests, and Benjamini-Hochberg False Discovery Rate (FDR) control to prove true statistical significance.
Mathematical Rigor & Calibration Methodology
1. The Calibration Gate
For every candidate verifier ( V ), we construct a confusion matrix over golden slices:
- False Pass Rate (FPR): How often does the verifier wrongly pass a broken output? $$\text{FPR} = \frac{\text{FP}}{\text{FP} + \text{TN}}$$
- Recall: How often does it correctly pass an authentic output? $$\text{Recall} = \frac{\text{TP}}{\text{TP} + \text{FN}}$$
A verifier is admitted if and only if: $$\text{FPR} \le 0.100 \quad \text{and} \quad \text{Recall} \ge 0.800$$
2. Wilson Score 95% Confidence Intervals
Standard asymptotic normal intervals collapse near 0% and 100% pass rates. DriftGate uses the continuity-corrected Wilson score interval for sample proportions:
$$w = \frac{\hat{p} + \frac{z^2}{2n} \pm z \sqrt{\frac{\hat{p}(1 - \hat{p})}{n} + \frac{z^2}{4n^2}}}{1 + \frac{z^2}{n}}$$
In our benchmark regression test (account_closure), the pass rate dropped from ( 40/40 ) (100%) to ( 0/40 ) (0%). DriftGate mathematically isolated this behavioral drift with:
$$\text{Effect Size} = -1.000, \quad \text{Wilson 95\% CI} = [-1.00, -0.85], \quad p = 1.872 \times 10^{-18}$$
How We Built It
DriftGate was engineered from the ground up as a production-grade infrastructure stack:
- Computational Core: FastAPI + Pydantic v2 for strictly validated agent trace schemas.
- Database & Vector Store: Neon Serverless Postgres with
pgvectorfor instant branching and 1536-dimensional cosine similarity indexing. - Agentic Orchestration: LangGraph 4-node state machine (
analyze_invariants$\rightarrow$draft_candidates$\rightarrow$self_critique$\rightarrow$emit). - Statistical Engine: SciPy, NumPy, and Scikit-Learn implementing HDBSCAN density clustering, Fisher's exact tests, two-proportion z-tests, and Benjamini-Hochberg multi-test corrections.
- Inference Gateway: Featherless AI providing OpenAI-compatible access across 32,000+ open-weights models (Llama 3.1/3.2/3.3, Qwen 2.5, RWKV v6) with SHA-256 response caching.
- Cinematic Frontend & 3D Observatory: Three.js, GSAP, and Vercel with a Pinterest-editorial aesthetic, interactive 3D WebGL node constellation, and a 30-Second Judge Mode.
- Deployment: Backend Web Service on Render with CORS middleware, Frontend on Vercel.
Challenges We Ran Into
- The LLM Judge Paradox: When we tested uncalibrated LLM judges (prompting a model with "Does this look correct?"), they passed the broken
"description"schema with an FPR of 1.000 (100% failure to detect bugs). They rubber-stamped broken outputs because the text "sounded polite." Building the Calibration Gate to mathematically detect and purge these uncalibrated judges was our biggest architectural breakthrough. - Boundary Stability in High-Dimensional Embeddings: Standard k-means clustering arbitrarily splits semantic spaces. We tuned HDBSCAN with custom min-cluster metrics on cosine distance matrices to discover true operational clusters without hardcoding domain schemas.
- Statistical Validity at Low Sample Counts: Naive evals report percentages based on 5 samples. We implemented dynamic test selection: switching to Fisher's Exact Test when ( n < 30 ) and Two-Proportion Z-Test when ( n \ge 30 ), coupled with Benjamini-Hochberg FDR control to prevent false alarms across multiple clusters.
Accomplishments That We're Proud Of
- 100% Test Coverage: 229 passed unit and integration tests verifying statistical math, database migrations, LangGraph state transitions, and HTML rendering.
- Zero-Vibes Infrastructure: Built an evaluation tool where the user never has to "trust" the evaluator blindly; every verifier publishes an auditable Calibration Card with verified FPR and Recall.
- The Featherless Pareto Frontier: DriftGate calculates the cross-model cost/accuracy Pareto curve across open-source models, identifying the cheapest model that passes $\ge 80\%$ of your suite.
- 30-Second Cinematic Judge Mode: Created a full Three.js spatial experience allowing hackathon judges and developers to understand the entire 7-stage engine in under 30 seconds.
What We Learned
- Vibes don't scale; invariants do. Deterministic Python AST checks and JSON schema validators consistently outperform uncalibrated LLM-as-a-judge prompts in both latency and accuracy.
- Regression testing for AI is fundamentally a statistical problem, not a generative problem. Confidence intervals and false discovery control are the missing layer in modern LLM ops.
What's Next for DriftGate
- GitHub Actions CI/CD Gate: A drop-in
driftgate-actionthat blocks pull requests if prompt changes cause a statistically significant behavioral drop (( p < 0.01 )). - Automated Prompt Repair: Feeding failing invariants back into LangGraph to automatically suggest the prompt diff required to restore baseline behavior.
- Real-Time OpenTelemetry Ingestion: Direct streaming connectors for LangSmith, Arize Phoenix, and OpenLLMetry.
Built With
- ai-infrastructure
- alembic
- docker
- fastapi
- featherless-ai
- gsap
- hdbscan
- html5
- jinja
- langgraph
- llmops
- machine-learning
- neon-postgres
- numpy
- pgvector
- prompt-engineering
- pydantic
- pytest
- python
- render
- scikit-learn
- scipy
- statistical-analysis
- three.js
- vercel
Log in or sign up for Devpost to join the conversation.