Inspiration
E-commerce businesses face a double-edged sword when deploying AI customer support: high recurring cloud API costs (OpenAI/Anthropic) and severe data privacy risks. To make matters worse, standard LLMs often hallucinate—making up fake coupon codes, wrong store policies, or false delivery dates that create customer service nightmares.
We wanted to build a zero-cost-per-token, 100% private, and completely offline solution. Our goal was to create a production-grade AI agent that lives entirely on the merchant's local machine, respects data privacy, and is engineered from the ground up to never lie to a customer.
What it does
Smart Customer Support Bot is an all-in-one, local e-commerce support ecosystem.
RAG-Powered Chat: Customers get instant, context-aware answers pulled straight from the store's internal FAQ knowledge base.
Anti-Hallucination Guardrails: If a question falls outside the known FAQs, the bot refuses to guess. Instead, it triggers a seamless human-handoff fallback.
Order Tracking: A public-facing portal where customers can view a visual, step-by-step progress stepper of their delivery status.
Full Admin Suite: A secure dashboard for store managers to handle analytics, perform CRUD operations on products and FAQs, view chat logs, and trigger a one-click reindexing of the vector store.
How we built it
We chose a modular, production-ready tech stack designed to optimize performance on consumer-grade hardware (even on just 8 GB of RAM):
The AI Core: We utilized Ollama running qwen2.5:3b for ultra-fast, lightweight text generation, alongside nomic-embed-text for creating high-quality semantic vector embeddings.
Vector & Storage: ChromaDB acts as our persistent local vector store handles cosine similarity matching, while SQLite managed via SQLAlchemy 2.0 drives our structured relational data.
The Backend: Built with FastAPI using a strict Clean Architecture pattern (Routers → Services → Repositories → Models) ensuring decoupling, high testability with PyTest, and automatic Swagger documentation.
The Frontend: A sleek, fully responsive UI built with Next.js 15 (App Router), TypeScript, Tailwind CSS, and Framer Motion for fluid animations and a modern dark mode theme.
Infrastructure: Orchestrated entirely via Docker Compose to handle volume caching for models, ensuring a true one-command deployment.
Challenges we ran into
Local Hardware Optimization: Running an LLM and an embedding model simultaneously on basic hardware can cause massive latency or out-of-memory errors. We solved this by carefully benchmarking lightweight models, landing on qwen2.5:3b, which balanced extreme speed with highly accurate instruction following.
Preventing Hallucinations Locally: Smaller models love to please the user, which leads to fabricating information. We spent hours refining a hyper-strict grounding system prompt combined with a strict cosine similarity threshold. If ChromaDB doesn't return context with a high confidence score, the backend bypasses generation entirely and flags a human handoff.
Docker/Ollama Race Conditions: During the initial setup, the backend would spin up and attempt to index data before Ollama had finished pulling the heavy 4GB model weights. We engineered a robust python startup script that pings Ollama’s health endpoints and safely waits for the models to be fully downloaded and cached before initiating migrations and data seeding.
Accomplishments that we're proud of
Zero Cloud Dependence: We successfully created an intelligent, context-aware AI application that requires zero API keys, zero internet connections (after the initial pull), and costs $0.00 to run endlessly.
True Clean Architecture: Even under hackathon time constraints, we didn't cut corners on code quality. The backend features a separated repository pattern, strict Pydantic v2 schemas, JWT authentication, and a mocked testing suite.
Seamless UX/DX: A developer can clone this repo, run docker compose up, and walk away. The app takes care of fetching models, seeding dummy store data, embedding documents, and delivering an immediately usable app at localhost:3000.
What we learned
We deepened our understanding of the RAG lifecycle—specifically the fine margins of word-bounded chunking and overlapping to ensure the LLM gets the cleanest context possible.
We discovered just how powerful open-source, local small language models (SLMs) have become. With proper prompt engineering, a 3B model can easily match or outperform massive commercial models on specialized, narrow domain tasks.
We mastered orchestrating multi-container Docker applications that rely on heavy, external hardware-dependent runtimes like Ollama.
What's next for Smart Customer Support Bot
Streaming Responses: Implementing Server-Sent Events (SSE) or WebSockets to stream the LLM's responses token-by-token to the frontend chat UI for a more natural user experience.
Live Agent Dashboard: Expanding the admin panel into a real-time socket-based inbox where human customer support reps can instantly view and take over conversations that were handed off by the bot.
Enterprise Scaling: Swapping the SQLite backend for PostgreSQL and migrating ChromaDB to pgvector to support multi-tenant workspaces and handle millions of products and FAQs seamlessly.
Built With
- chromadb
- fastapi
- next.js
- ollama
- python
- qwen2.5
- sqlalchemy
- sqlite
- tailwind
- typescript


Log in or sign up for Devpost to join the conversation.