Inspiration

Shopping search usually assumes that users already know exactly what they want.

Real shoppers do not behave that way.

They start with vague requests, reveal preferences gradually, reject options, and sometimes change their minds entirely.

A good store clerk does not expect shoppers to arrive with a perfectly specified query. They ask questions, remember preferences, and adapt as those preferences change.

We wanted to bring that behavior to online shopping.

That is why we built Shopping Clerk.

How we framed the problem

We framed conversational shopping as a sequential search problem under incomplete and evolving intent, rather than a one-shot recommendation task.

At every turn, the system must solve three problems:

  1. Track what is currently true
    Which preferences are active, uncertain, or superseded?

  2. Retrieve and rank with incomplete evidence
    Preserve enough candidate recall while using known constraints to move relevant products toward the top.

  3. Decide when to ask
    When intent is still ambiguous, a useful clarification can change the next retrieval and ranking step.

This framing led us to separate state tracking, retrieval, ranking, and clarification into a closed loop instead of treating the entire conversation as one long search query.

What it does

Shopping Clerk is a multi-turn conversational search and recommendation system for a 50,000-product fashion catalog.

A shopper can begin with something as simple as:

"I need a jacket."

Instead of forcing them to formulate the perfect query upfront, Shopping Clerk can ask useful clarification questions, remember their answers, retrieve new candidates, rerank recommendations, and update its understanding as the conversation evolves.

The system also handles changes of mind.

For example, if the shopper first asks for a black leather jacket and later says:

"Actually, make it denim."

Shopping Clerk supersedes the leather preference while preserving unrelated constraints such as "black" and "jacket." The recommendations then update using the revised intent.

How we built it

We designed Shopping Clerk as a closed-loop pipeline:

Conversation → Living State → Query Construction → Multi-Route Retrieval → Constraint Ranking → Free-Text Reranking → Clarification → Recommendations

Living State

Instead of interpreting every message as an isolated query, the system maintains a structured representation of the shopper's current intent.

Preferences can remain active or become superseded when the shopper changes their mind.

This allows:

BLACK + JACKET + LEATHER

to become:

BLACK + JACKET + DENIM

without losing unrelated preferences.

Multi-Route Retrieval

The current state is converted into a retrieval query and sent through multiple catalog routes.

The shipped system combines BM25 full-text retrieval and category-based retrieval, then merges their results into a shared candidate pool.

This separation is deliberate: retrieval is responsible for making relevant products available to rank, rather than forcing the first search stage to solve the entire recommendation problem.

Constraint Ranking

Candidates are ranked using known product constraints such as category, material, color, size, style, brand, budget, and other available attributes.

Because real product catalogs contain incomplete metadata, we explicitly distinguish:

MATCH ≠ UNKNOWN ≠ MISMATCH

A missing value therefore does not automatically become evidence that a product violates the shopper's request.

Free-Text Reranking

Structured attributes cannot capture everything a shopper says.

After constraint ranking, a lexical free-text reranker uses additional evidence from the shopper's own words to improve the ordering of the candidate window.

This gives the system another opportunity to surface relevant products without discarding the structured constraints already learned from the conversation.

Clarification

Clarification is part of the search loop rather than a separate chatbot feature.

When intent remains ambiguous, Shopping Clerk can ask for another useful attribute and feed the answer directly back into state, retrieval, and ranking.

The resulting interaction becomes:

Ask → Learn → Retrieve → Rank → Adapt

Multi-turn adaptation

The key idea behind Shopping Clerk is that conversation does not simply provide more text — it changes the search state itself.

When the shopper reveals a new preference, the next search uses that new evidence.

When a preference changes, only the conflicting evidence is superseded while unrelated preferences remain active.

The result is a search experience that evolves with the shopper instead of restarting from scratch after every message.

Technical analysis

We evaluated Shopping Clerk not only by its final score, but also by decomposing where failures occurred.

The official evaluator measures three complementary objectives:

  • HR@10 — whether the correct product reaches the Top 10
  • MRR — how highly the correct product is ranked
  • MTTC — how quickly the system converges to it across the conversation

Our final frozen system achieved:

  • HR@10: 0.9000
  • MRR: 0.591530
  • MTTC: 4.390
  • Technical Score: 0.759659

At the retrieval level, scoring-eligible Candidate Recall@300 reached 0.9750, meaning the target entered the candidate pool in 195 of 200 evaluation sessions.

Only 5 sessions remained retrieval failures. Of the remaining misses, 15 already contained the target in the candidate pool but still failed to place it in the Top 10.

That failure decomposition changed how we optimized the system: once candidate recall reached 97.5%, the main remaining problem was no longer finding more products — it was ranking the products we already had more effectively.

We also used staged ablations rather than assuming every feature we built should ship.

Components were enabled and measured incrementally, and changes that did not establish sufficient evidence were reverted rather than kept simply because they added complexity.

For the final configuration, we also tested how many candidates the reranker should be allowed to see.

Increasing the reranking window from 50 to 200 candidates improved Technical Score from 0.722979 to 0.759659, with 10 sessions gained and 0 lost. The improvement was supported by both hit-set and paired score tests, every evaluation scenario improved or held, and the reranker remained well inside its runtime budget with zero timeouts or fallbacks.

This gave us a final configuration based on measured contribution, failure decomposition, significance testing, and regression checks rather than feature count or intuition.

Why it matters

Shopping Clerk is designed around a simple user problem: shoppers often do not know the perfect query before they start searching.

Traditional search puts the burden on the shopper to repeatedly translate evolving preferences into new keyword queries.

Shopping Clerk shifts more of that burden to the system.

By asking for useful missing information, remembering what has already been learned, and adapting when preferences change, the system aims to reduce repeated searching and help shoppers converge on relevant products in fewer conversational turns.

The same interaction pattern can extend beyond fashion to domains such as electronics, furniture, beauty, and other product categories where shoppers often discover what they want while they browse.

Built for practical deployment

We deliberately kept the shipped architecture lightweight and reproducible.

The retrieval stack operates over the 50,000-product catalog using SQLite FTS5 and BM25, with a bounded candidate pool and deterministic downstream processing.

The Free-Text Reranker is lexical and operates directly over the candidate window, keeping the runtime path compact.

In the final configuration, the reranker produced zero timeouts and zero fallbacks during evaluation.

The complete system also passed 615 tests on both Python 3.9.6 and Python 3.12.2, with evaluation results reproduced identically across multiple Python hash seeds.

This keeps the current architecture grounded, reproducible, and practical to run without requiring heavyweight external infrastructure.

Challenges

Sparse product metadata

The catalog does not provide complete structured attributes for every product.

Treating missing values as hard mismatches would incorrectly remove potentially relevant products, so we explicitly distinguish UNKNOWN from MISMATCH.

Evolving intent

A shopper changing one preference should not erase everything learned earlier.

We therefore track active and superseded evidence separately so that a request such as "make it denim" replaces the material preference without losing unrelated constraints such as category or color.

Knowing when to ask

Clarification is useful only when the answer is likely to improve the search.

We therefore treat clarification as part of the retrieval loop rather than asking questions after every turn.

Avoiding benchmark-only optimization

A higher benchmark score does not automatically mean a better shopping experience.

Some configurations produced attractive point estimates on the evaluator but relied on behavior we did not want the final system to depend on.

Rather than selecting every public-score maximum, we used regression checks, ablations, failure analysis, and statistical evidence to decide what actually shipped.

What we learned

One of our biggest lessons was that search quality depends not only on the retrieval algorithm, but also on the information available to it and on what happens after retrieval.

By the final pipeline, scoring-eligible candidate recall reached 97.5%. At that point, retrieval failures had fallen to only 5 sessions, while the larger remaining failure class already contained the target in the candidate pool.

That shifted our engineering focus downstream: from simply retrieving more products to ranking the candidates we already had more effectively.

We also learned that a feature's value cannot be judged in isolation. The same component can behave differently depending on which other stages are already active, so we evaluated changes as part of the pipeline they actually operate in.

Most importantly, we treated measurement as part of the product design, rather than as an after-the-fact benchmark.

Built With

Share this project:

Updates

Submission history