Inspiration
Picture a finance team getting an email at 4 PM: tomorrow, the CFO is meeting the sales team and a client, and needs to walk through last quarter's revenue, flag anomalies, and show where the business stands. Normally that's a scramble pull the data, write the SQL, build a dashboard, sanity-check every number two to five hours of work, the night before a meeting that was supposed to be routine.
We wanted to replace that scramble with a link. Open it, ask a question in plain English, get a real answer backed by the actual warehouse not a chatbot that sounds confident, one that's actually grounded in data and honest about what it doesn't know.
What it does
It's a multi-agent AI copilot for retail analytics. Ask it "What was the total revenue last month?" or "Which customer segment drives the most revenue?" and it routes your question through a 12-node LangGraph pipeline: classifies intent, retrieves relevant business context, writes and validates real SQL, executes it against a live PostgreSQL warehouse, and critically cross-checks its own reasoning against the warehouse before answering. If the knowledge base disagrees with the live data, the warehouse wins, every time.
It also forecasts revenue with a trained XGBoost model, flags real anomalies with Isolation Forest, and when it genuinely isn't confident in an answer logs it to a human review queue instead of returning a low-quality answer with false confidence.
How I built it
- Backend: FastAPI, orchestrating a LangGraph multi-agent pipeline (router → RAG retrieval → context compression → SQL generation/validation/execution → reasoning → verification → recommendation → critic).
- Data layer: PostgreSQL on Neon, holding 1M+ real UK retail transactions, with 15 analytics views built on top.
- RAG layer: ChromaDB with a local ONNX Runtime embedder no PyTorch dependency in production, after a real optimization journey (see Challenges).
- LLM layer: Groq as the primary model, with a genuine Gemini fallback and an in-process circuit breaker for when Groq's free tier throttles under load.
- Frontend: Next.js, Tailwind, Framer Motion streams the agent's response token-by-token, with generative UI for charts when the data calls for it.
- Deployment: Docker container on Azure App Service (backend), Vercel (frontend), Neon (database).
Challenges I ran into
The fallback that wasn't. Our README claimed a "Gemini fallback" for months. It was dead code never imported, never called. We only found out because a live evaluation run crashed mid-question when Groq rate-limited it. Fixing it for real took four attempts: one experimental Gemini model with restrictive rate limits, one model that didn't exist for our API key, one that had been deprecated (per Google's own error message, which told us the replacement), and finally one that worked confirmed with a standalone test, then verified in production logs.
The six-day outage that wasn't. Our uptime monitor reported the API down for six straight days. It never actually was every real request succeeded the whole time. The monitor was sending HEAD requests to a route only defined for GET, so every check got a 405 and was read as "down." A one-line route fix resolved it, confirmed the moment it deployed.
The Docker size problem. We went through sentence-transformers → Gemini embedding API → HuggingFace API → local ONNX runtime while chasing out-of-memory errors on a constrained deployment tier. Verified by building the "before" and "after" versions from actual git history: 14.4GB down to 2.75GB.
Accomplishments that I am proud of
The system caught a real bug in its own knowledge base a wrong product name that had been sitting there since day one by cross-checking against the live warehouse and correcting itself.
- A genuinely tested Human-in-the-Loop pipeline: built, tested end-to-end (write → read → resolve, through the real API, against the real database), and confirmed working in production.
- Every number in our documentation is something we personally measured, not estimated including the uncomfortable ones, like a RAGAS faithfulness score of 0.38 that we reported honestly instead of hiding.
What I learned
Production AI systems fail in ways that are invisible until you actually load-test them. Our fallback logic looked complete in code review; it took a live rate-limit event to reveal it did nothing. Our uptime monitor looked broken; it took inspecting the actual HTTP method to find a one-line fix. The lesson that stuck: verify against the real system, not against what the code appears to do on paper.
What's next
- A statistically robust RAGAS benchmark (current numbers are from a single run and stated as such).
- A distributed circuit breaker (Redis-backed) for when this scales beyond one container instance.
- Root-causing the 4.7–7s Neon query latency we found but haven't yet fixed.
- A small review UI on top of the existing
/review/pendingand/review/{id}/resolveendpoints.
Built With
- azure
- azure-app-service
- chromadb
- docker
- eda
- fastapi
- framer-motion
- gemini-api
- groq
- hitl
- langgraph
- mlflow
- neon
- nextjs
- onnx-runtime
- pandas
- postgresql
- python
- ragas
- scikit-learn
- sqlalchemy
- tableau
- tailwindcss
- vercel
- xgboost
Log in or sign up for Devpost to join the conversation.