Intent-Aware Conversational Product Search
Inspiration
Shopping search becomes difficult when customers cannot describe everything they want in one query. Someone may begin with “I’m looking for a wallet,” later reveal that leather is essential, and then replace an earlier color preference. A useful shopping assistant must do more than match keywords: it must remember confirmed requirements, recognize intent changes, and ask questions that meaningfully narrow the options.
For TikTok TechJam Track 4, we built a conversational recommender that finds and ranks a customer’s hidden target product within ten turns. We focused on accuracy, explainability, reproducibility, and offline operation rather than relying on an external LLM or network service.
What it does
Our agent searches a frozen catalog of 50,000 clothing, shoes, jewelry, and accessory products. On every turn, it:
- identifies whether the customer is browsing or ready to buy
- extracts preferences such as category, material, color, size, style, use case, feature, and budget
- updates conversation memory and removes preferences that have been overridden
- retrieves and reranks relevant catalog products
- asks a useful clarification question when more information could improve the result
The system handles all four challenge scenarios:
- Buying: combines the category with explicit must-have requirements
- Browsing: supports broader product discovery when the request is vague
- Intent Override: removes superseded preferences while preserving compatible constraints
- Boundary: continues gracefully when the customer has no preference for an attribute
This directly addresses the challenge metrics: multi-route retrieval protects Hit Rate@10, field-aware ranking improves reciprocal rank, and high-information questions reduce the number of turns required to find the target.
How we built it
Structured conversation state
Each session stores typed preference slots with their value, source, turn, and active status. This lets the agent distinguish a confirmed requirement from a soft preference and selectively retire outdated information after an intent change.
The state also records the current intent, conversation phase, recently asked questions, declined attributes, and a bounded history of recent turns.
Intent-aware retrieval
The catalog is loaded into an in-memory SQLite FTS5 index with Porter stemming. The agent uses several complementary retrieval routes:
- a precise AND query over the category and strongest constraints
- independent searches for each constraint to preserve recall
- an OR fallback when a restrictive query returns no candidates
- a lightweight feature-hash discovery route for sparse browsing results
The discovery route uses a deterministic 192-dimensional NumPy representation built from word and character fragments. It is not a pretrained language model, but it provides an inexpensive local fallback for partial words and ordinary morphology.
Field-aware ranking
Candidates are ranked using:
- coverage of every active requirement
- exact phrase and structured-field evidence
- canonical category-tail matching
- greater weight for recently confirmed constraints
- a modest rating and review-count tie-breaker
This prevents broad category mentions or metadata labels from receiving the same credit as exact product evidence.
Adaptive clarification
The agent estimates which question is most likely to reduce uncertainty among the strongest candidates. For an attribute (a), we approximate: $$ \operatorname{utility}(a) = \operatorname{answerability}(a) \left( \operatorname{coverage}(a) - \sum_v P(v \mid a)^2 \right) $$
Known, previously asked, or explicitly declined attributes are penalized. Early turns collect broad evidence; later questions focus on the remaining ambiguity.
Candidate diagnostics
To avoid tuning blindly against a single aggregate score, we built a local candidate diagnostics tool for the released development set. The tool replays sessions through the evaluator and records whether slow or unresolved cases are mainly retrieval, ranking, dialog, or intent-override problems.
For each session, it can inspect the candidate pool, the target's rank inside that pool, the returned recommendation rank, active constraints, current workflow strategy, and next-question utility. It also produces aggregate failure counts, scenario breakdowns, rank distributions, and an "improvement_focus" summary that identifies the highest-value sessions to inspect first.
Results
On all 200 released development sessions, the frozen submission achieved:
| Metric | Result |
|---|---|
| Hit Rate@10 | 1.000000 |
| Mean Reciprocal Rank | 0.967464 |
| Mean Turns to Conversion | 2.090 |
| Technical score | 0.968439 |
| Model/API tokens | 0 |
For an additional robustness check, we evaluated on 1,000 locally generated target products excluded from the public target set. The system achieved 0.950 Hit Rate@10, 0.885 MRR, and a 0.902 technical score. This was a synthetic stress test using the released protocol—not organizer-private data.
No sample IDs, target IDs, or public labels are read by the production agent.
Development tools used
- Python 3.12
- Git for version control and experiment tracking
- Python
unittestfor dialog, retrieval, evaluator, and submission-package tests - PowerShell and POSIX command-line tools for evaluation, benchmarking, and packaging
The project is editor-agnostic and does not require Colab or Jupyter.
APIs used
No external APIs are used. The agent requires no API keys, makes no network requests, and consumes no model tokens. The complete judging path works offline rather than relying on a reduced fallback mode.
Libraries and frameworks used
- Python standard library, including
sqlite3,re,json,pathlib,math,zlib, anddataclasses - NumPy 2.3.5 for feature-hash vector operations
- SQLite FTS5 with Porter stemming for weighted lexical retrieval
We intentionally avoided hosted LLMs, PyTorch, TensorFlow, and external vector databases to keep the submission deterministic and CPU-friendly.
Datasets and assets used
- Amazon Reviews 2023 — Clothing, Shoes and Jewelry, from McAuley Lab at UCSD
- The challenge-provided frozen catalog of 50,000 products
- 200 released development sessions across Buying, Browsing, Intent Override, and Boundary scenarios
- A locally generated 1,000-target stress test containing products excluded from the public targets
- The challenge-provided evaluator, API contract, scoring configuration, and BM25 baseline
The agent uses only participant-visible catalog metadata and the aggregate profile supplied at inference. It does not use raw user identities, private evaluation data, timestamps, review text, or hidden intent cards.
Challenges we faced
Balancing recall and rank
Returning ten products immediately can improve Hit Rate@10 but record a poor reciprocal rank before the customer reveals distinguishing evidence. Returning too few products can increase misses. Balancing those objectives required careful conversation and recommendation timing.
Preserving category and requirements
An early version accidentally replaced the category when a Buying message also contained a hard requirement. Searching for “leather” across the entire catalog was much weaker than searching for “wallets + leather.” Storing both as independent slots produced a major precision improvement.
Handling intent changes
Clearing every constraint after an override loses valuable information, while retaining everything allows stale preferences to contaminate the new search. Typed slots allowed us to retire only the superseded preference.
Separating evidence from metadata noise
Flattened catalog text can make ancestor categories, field labels, and exact product attributes appear equally meaningful. Canonical category matching and field-level evidence helped us distinguish genuine compatibility from incidental text overlap.
What we learned
The biggest lesson was that conversational retrieval is as much a state-management problem as a ranking problem. Even a strong retriever performs poorly if the system loses a category, keeps a retired preference, or asks an unhelpful question.
We also learned that:
- candidate recall and final ordering should be evaluated separately
- structured field evidence can be more valuable than adding another broad retrieval route
- clarification should depend on the current candidate distribution
- deterministic systems can be competitive when the catalog schema is rich
- public-set improvements must be validated against disjoint target products
What we are proud of
- Improved the released BM25 baseline technical score from approximately 0.125 to 0.968
- Achieved perfect public Hit Rate@10 with zero API cost and zero model tokens
- Preserved strong ranking performance on a disjoint 1,000-target stress test
- Built explicit support for intent evolution and customer boundary responses
- Delivered a reproducible, no-network submission with pinned dependencies, manifest hashing, and strict interface validation
What is next
Next, we would add uncertainty-driven recall slots, replace fixed recommendation width with confidence-calibrated Top-1/Top-3/Top-10 decisions, strengthen parsing for less structured and multilingual requests, and reduce cold-start cost through lazy or cached discovery indexing.
Longer term, an offline sentence-embedding model or controlled query rewriter could improve semantic recall while retaining the deterministic lexical pipeline as a reliable fallback.
Log in or sign up for Devpost to join the conversation.