Intent-Aware Conversational Product Search

Inspiration

Shopping search becomes difficult when customers cannot describe everything they want in one query. Someone may begin with “I’m looking for a wallet,” later reveal that leather is essential, and then replace an earlier color preference. A useful shopping assistant must do more than match keywords: it must remember confirmed requirements, recognize intent changes, and ask questions that meaningfully narrow the options.

For TikTok TechJam Track 4, we built a conversational recommender that finds and ranks a customer’s hidden target product within ten turns. We focused on accuracy, explainability, reproducibility, and offline operation rather than relying on an external LLM or network service.

What it does

Our agent searches a frozen catalog of 50,000 clothing, shoes, jewelry, and accessory products. On every turn, it:

  1. identifies whether the customer is browsing or ready to buy
  2. extracts preferences such as category, material, color, size, style, use case, feature, and budget
  3. updates conversation memory and removes preferences that have been overridden
  4. retrieves and reranks relevant catalog products
  5. asks a useful clarification question when more information could improve the result

The system handles all four challenge scenarios:

  • Buying: combines the category with explicit must-have requirements
  • Browsing: supports broader product discovery when the request is vague
  • Intent Override: removes superseded preferences while preserving compatible constraints
  • Boundary: continues gracefully when the customer has no preference for an attribute

This directly addresses the challenge metrics: multi-route retrieval protects Hit Rate@10, field-aware ranking improves reciprocal rank, and high-information questions reduce the number of turns required to find the target.

How we built it

Structured conversation state

Each session stores typed preference slots with their value, source, turn, and active status. This lets the agent distinguish a confirmed requirement from a soft preference and selectively retire outdated information after an intent change.

The state also records the current intent, conversation phase, recently asked questions, declined attributes, and a bounded history of recent turns.

Intent-aware retrieval

The catalog is loaded into an in-memory SQLite FTS5 index with Porter stemming. The agent uses several complementary retrieval routes:

  • a precise AND query over the category and strongest constraints
  • independent searches for each constraint to preserve recall
  • an OR fallback when a restrictive query returns no candidates
  • a lightweight feature-hash discovery route for sparse browsing results

The discovery route uses a deterministic 192-dimensional NumPy representation built from word and character fragments. It is not a pretrained language model, but it provides an inexpensive local fallback for partial words and ordinary morphology.

Field-aware ranking

Candidates are ranked using:

  • coverage of every active requirement
  • exact phrase and structured-field evidence
  • canonical category-tail matching
  • greater weight for recently confirmed constraints
  • a modest rating and review-count tie-breaker

This prevents broad category mentions or metadata labels from receiving the same credit as exact product evidence.

Adaptive clarification

The agent estimates which question is most likely to reduce uncertainty among the strongest candidates. For an attribute (a), we approximate: $$ \operatorname{utility}(a) = \operatorname{answerability}(a) \left( \operatorname{coverage}(a) - \sum_v P(v \mid a)^2 \right) $$

Known, previously asked, or explicitly declined attributes are penalized. Early turns collect broad evidence; later questions focus on the remaining ambiguity.

Candidate diagnostics

To avoid tuning blindly against a single aggregate score, we built a local candidate diagnostics tool for the released development set. The tool replays sessions through the evaluator and records whether slow or unresolved cases are mainly retrieval, ranking, dialog, or intent-override problems.

For each session, it can inspect the candidate pool, the target's rank inside that pool, the returned recommendation rank, active constraints, current workflow strategy, and next-question utility. It also produces aggregate failure counts, scenario breakdowns, rank distributions, and an "improvement_focus" summary that identifies the highest-value sessions to inspect first.

Results

On all 200 released development sessions, the frozen submission achieved:

Metric Result
Hit Rate@10 1.000000
Mean Reciprocal Rank 0.967464
Mean Turns to Conversion 2.090
Technical score 0.968439
Model/API tokens 0

For an additional robustness check, we evaluated on 1,000 locally generated target products excluded from the public target set. The system achieved 0.950 Hit Rate@10, 0.885 MRR, and a 0.902 technical score. This was a synthetic stress test using the released protocol—not organizer-private data.

No sample IDs, target IDs, or public labels are read by the production agent.

Development tools used

  • Python 3.12
  • Git for version control and experiment tracking
  • Python unittest for dialog, retrieval, evaluator, and submission-package tests
  • PowerShell and POSIX command-line tools for evaluation, benchmarking, and packaging

The project is editor-agnostic and does not require Colab or Jupyter.

APIs used

No external APIs are used. The agent requires no API keys, makes no network requests, and consumes no model tokens. The complete judging path works offline rather than relying on a reduced fallback mode.

Libraries and frameworks used

  • Python standard library, including sqlite3, re, json, pathlib, math, zlib, and dataclasses
  • NumPy 2.3.5 for feature-hash vector operations
  • SQLite FTS5 with Porter stemming for weighted lexical retrieval

We intentionally avoided hosted LLMs, PyTorch, TensorFlow, and external vector databases to keep the submission deterministic and CPU-friendly.

Datasets and assets used

  • Amazon Reviews 2023 — Clothing, Shoes and Jewelry, from McAuley Lab at UCSD
  • The challenge-provided frozen catalog of 50,000 products
  • 200 released development sessions across Buying, Browsing, Intent Override, and Boundary scenarios
  • A locally generated 1,000-target stress test containing products excluded from the public targets
  • The challenge-provided evaluator, API contract, scoring configuration, and BM25 baseline

The agent uses only participant-visible catalog metadata and the aggregate profile supplied at inference. It does not use raw user identities, private evaluation data, timestamps, review text, or hidden intent cards.

Challenges we faced

Balancing recall and rank

Returning ten products immediately can improve Hit Rate@10 but record a poor reciprocal rank before the customer reveals distinguishing evidence. Returning too few products can increase misses. Balancing those objectives required careful conversation and recommendation timing.

Preserving category and requirements

An early version accidentally replaced the category when a Buying message also contained a hard requirement. Searching for “leather” across the entire catalog was much weaker than searching for “wallets + leather.” Storing both as independent slots produced a major precision improvement.

Handling intent changes

Clearing every constraint after an override loses valuable information, while retaining everything allows stale preferences to contaminate the new search. Typed slots allowed us to retire only the superseded preference.

Separating evidence from metadata noise

Flattened catalog text can make ancestor categories, field labels, and exact product attributes appear equally meaningful. Canonical category matching and field-level evidence helped us distinguish genuine compatibility from incidental text overlap.

What we learned

The biggest lesson was that conversational retrieval is as much a state-management problem as a ranking problem. Even a strong retriever performs poorly if the system loses a category, keeps a retired preference, or asks an unhelpful question.

We also learned that:

  • candidate recall and final ordering should be evaluated separately
  • structured field evidence can be more valuable than adding another broad retrieval route
  • clarification should depend on the current candidate distribution
  • deterministic systems can be competitive when the catalog schema is rich
  • public-set improvements must be validated against disjoint target products

What we are proud of

  • Improved the released BM25 baseline technical score from approximately 0.125 to 0.968
  • Achieved perfect public Hit Rate@10 with zero API cost and zero model tokens
  • Preserved strong ranking performance on a disjoint 1,000-target stress test
  • Built explicit support for intent evolution and customer boundary responses
  • Delivered a reproducible, no-network submission with pinned dependencies, manifest hashing, and strict interface validation

What is next

Next, we would add uncertainty-driven recall slots, replace fixed recommendation width with confidence-calibrated Top-1/Top-3/Top-10 decisions, strengthen parsing for less structured and multilingual requests, and reduce cold-start cost through lazy or cached discovery indexing.

Longer term, an offline sentence-embedding model or controlled query rewriter could improve semantic recall while retaining the deterministic lexical pipeline as a reliable fallback.

Built With

Share this project:

Updates

Submission history