About the Project
Inspiration
Online shopping often begins with an incomplete idea: “I need comfortable shoes,” “something for a wedding,” or “a gift under $100.” Traditional search engines expect shoppers to know the right keywords and manually compare hundreds of products.
We built ShopCopilot to make product discovery feel more like speaking with a helpful store assistant. Instead of repeatedly showing broad results, the agent asks useful questions, remembers the shopper’s answers, and narrows the search with every interaction.
Our guiding idea was simple:
The best shopping agent should maximize the value of every conversation turn.
What It Does
ShopCopilot is a conversational shopping agent built for Track 4. It searches a catalog of 50,000 products and attempts to identify the shopper’s target product within ten turns.
The agent can:
- Distinguish between high-intent Buying requests and exploratory Browsing requests
- Extract requirements such as category, colour, material, size, brand, style and budget
- Ask clarification questions that reduce the candidate space
- Remember preferences throughout a conversation
- Replace outdated preferences when the shopper changes their mind
- Recognize “no preference” answers and avoid asking about that attribute again
- Return ranked recommendations on every turn, including turns where it asks a question
- Use profile information only as a soft signal, ensuring the shopper’s current request always takes priority
How We Built It
ShopCopilot uses a local, deterministic retrieval and ranking pipeline.
First, the agent converts the conversation into three search views: the accumulated conversation, the latest message and the extracted constraints. It searches these views using field-weighted BM25 through SQLite FTS5 and lightweight hashed TF–IDF with cosine similarity.
We combine the resulting ranked lists using Reciprocal Rank Fusion. An in-memory Boolean inverted index then forms a high-recall “guarantee pool” by finding products that contain the shopper’s constraint tokens.
Candidates are reranked using:
- Explicit constraint violations
- IDF-weighted constraint coverage
- Retrieval relevance
- Title and feature overlap
- Field-specific matches
- Product popularity
- Inferred conversational context
- Optional personalization signals
For clarification, the agent analyzes the remaining products and chooses questions using candidate-set size, confidence, entropy and expected set reduction. This helps it ask about attributes that genuinely distinguish the remaining options.
The main technology stack includes:
- Python
- SQLite FTS5
- NumPy
- Pydantic
- Pytest
- JSON and JSONL
- Optional OpenRouter and Gemini integrations
The scored path remains fully local and does not require an external LLM or vector database.
Research and Experimentation
We studied established information-retrieval and conversational-search research while improving the system. Our implementation draws directly from BM25, TF-IDF, cosine similarity, feature hashing and Reciprocal Rank Fusion.
Research on conversational clarification also reinforced our decision to select questions using the current dialogue and retrieval state. Instead of reproducing a neural question-selection model, we implemented a deterministic, catalog-grounded approach using entropy and candidate-set reduction.
We also built an optional LLM reranking layer via OpenRouter (nvidia/nemotron-3-super-120b-a12b:free). However, our 4-case evaluation showed no improvement in Hit Rate (ΔHit 0) or MRR (ΔMRR −0.005) while adding 2.2–8.3s per call — 636% to 2,676% more than the local no-LLM p50 of 299 ms, ~7.4× to 27.8× slower. Based on those results, we disabled it in the default submission path. This kept ShopCopilot faster, cheaper, and more reproducible.
Challenges We Faced
One major challenge was balancing recall with ranking precision. Increasing the candidate pool rescued targets that earlier retrieval stages missed, but it also introduced more competing products that could displace the target from the Top 10.
Natural-language constraint extraction created several unexpected problems. Short letters could be mistaken for clothing sizes, unrelated numbers could be interpreted as budgets, and phrases such as “forget blue, I want red” required negation-aware handling.
Intent changes were another difficult area. Our original approach erased too much conversation history when the shopper changed a preference. Evaluation showed that preserving unrelated valid constraints while replacing only the contradicted field performed better.
Clarification also involved a trade-off: additional questions can improve accuracy, but unnecessary questions increase Mean Turns to Conversion (MTTC). We therefore added confidence and turn-budget gates to decide when another question is worth asking.
Finally, we had to avoid overfitting to the 200 public sessions. We created a reproducible development and holdout split, recorded rejected experiments, and used per-scenario diagnostics rather than relying only on one aggregate score.
What We Learned
We learned that strong conversational search depends on much more than selecting a powerful model.
Precise state management and constraint extraction often produced larger improvements than adding model complexity. We also learned that asking one useful question can be more valuable than retrieving hundreds of additional products.
Most importantly, we learned to let measurements determine the architecture. Features that sounded advanced were removed or disabled when they failed to improve the official metrics. Simpler methods remained when they were faster, safer and more reliable.
Results
On the official Track 4 public evaluation, ShopCopilot achieved a best result of:
- Hit Rate@10: 0.880
- MRR: 0.4916
- MTTC: 3.375
- TechnicalScore: 0.7400
The official weak starter achieved a TechnicalScore of 0.1067. ShopCopilot therefore achieved approximately 7 times the provided baseline, without requiring an external LLM or vector database in its scored path.
ShopCopilot demonstrates that an effective shopping copilot can be intelligent, explainable and practical without depending on expensive infrastructure.
Log in or sign up for Devpost to join the conversation.