Inspiration

Online shopping rarely begins with a perfectly specified query. A customer may start with “I need a lightweight camera for travel,” later add a budget, change a preferred brand, or say that an attribute no longer matters.

Traditional search systems usually treat each message as an isolated query. They forget earlier requirements, retain preferences that the user has already replaced, or retrieve and rank hundreds of products before realizing that the request is too broad.

We built Narrow to explore a different idea: conversational memory should not be passive chat history. It should be a validated, executable state that directly controls how the system searches, filters, ranks, and asks questions.

What it does

Narrow is a stateful conversational shopping copilot that turns a multi-turn conversation into an adaptive product-search policy.

For every user turn, Narrow:

  1. Interprets the latest message together with the maintained conversation state.
  2. Produces a structured StatePatch containing additions, replacements, removals, resets, and explicit “no preference” signals.
  3. Validates the patch before updating active and superseded constraints.
  4. Builds a canonical query that preserves all currently active requirements.
  5. Selects a dynamic retrieval policy for Buying, Browsing, or Unknown intent.
  6. Retrieves candidates through lexical, dense-semantic, and structured-attribute routes.
  7. Fuses the routes using policy-weighted Reciprocal Rank Fusion:

$$RRF(d)=\sum_r \frac{w_r}{k+\operatorname{rank}_r(d)}$$

  1. Applies hard constraints, soft preference boosts, and unknown-aware handling.
  2. Reranks the resulting candidates using learned ranking features.
  3. Either recommends the Top 10 products or asks one high-information clarification question.

Narrow also remembers which attributes have already been asked, which questions remain pending, which constraints were superseded, and which products were previously recommended.

An anonymized user profile can influence ranking through soft preference evidence, but it never overrides explicit requirements from the current session.

How we built it

The backend is implemented in Python as a LangGraph workflow.

DeepSeek performs intent understanding and dialogue decisions. Instead of allowing the LLM to directly mutate the search state, it returns a bounded Pydantic StatePatch. A deterministic validation layer normalizes values, rejects contradictions, and applies the patch to the canonical state.

The state contains:

  • Active and superseded constraints
  • Explicit “no preference” attributes
  • Asked attributes and pending questions
  • Question history
  • Previously recommended products
  • Semantic and lexical queries
  • Retrieval intent, policy, and diagnostics

Retrieval runs entirely in memory and combines three complementary routes:

  • BM25 lexical retrieval for exact models, brands, and keywords
  • Dense retrieval for semantic intent and use cases
  • Attribute retrieval for price, color, material, size, style, and features

The runtime controller changes each route’s weight, retrieval depth, and final candidate-pool size according to the current intent and constraint specificity.

After fusion, Narrow uses tri-state constraint evidence: MATCH, VIOLATE, or UNKNOWN. A product is removed only when the catalog provides explicit evidence that it violates a reliable hard constraint. Missing metadata is preserved as unknown instead of being incorrectly treated as false.

For final ranking, we built a LambdaMART pipeline using lexical, semantic, attribute, constraint, quality, novelty, and profile-match features. We also freeze the training-time IDF table with the model to avoid train–serve feature skew.

The demo interface is built with Vue, TypeScript, Vite, and Pinia. A local Starlette API exposes conversations, maintained state, retrieval plans, route diagnostics, candidates, ranking features, and node-level traces.

Challenges we ran into

Maintaining intent without carrying stale preferences

Multi-turn intent is not monotonic. Users can add requirements, replace one value, retract all soft preferences, or restart with an entirely new product category.

We addressed this with an explicit slot lifecycle. Constraints move between active and superseded states rather than disappearing silently. Category changes also reset incompatible question and recommendation history.

Incomplete catalog metadata

A missing material or price does not prove that a product violates the request. Treating missing values as violations caused relevant products to disappear during coarse ranking.

We introduced unknown-aware constraint evaluation so that only explicit contradictions trigger hard filtering.

Train–serve skew in reranking

Our earlier reranking weights were trained against a different coarse-ranking distribution. Once dynamic route weights changed the RRF score distribution, retraining the final ranker alone produced little improvement.

We rebuilt the training pipeline using candidates generated by the same dynamic retrieval stack used at runtime and froze the feature schema and global IDF data with the model.

LLM nondeterminism and evaluation integrity

Online LLM evaluations are not perfectly deterministic. We therefore retained every session and failed turn, disabled silent offline fallback during online evaluation, recorded model usage, and exported node-level traces for state, retrieval, fusion, ranking, and dialogue decisions.

Accomplishments that we're proud of

On a matched 200-session online evaluation using DeepSeek V4 Flash and the frozen LambdaMART bundle, Narrow achieved:

  • Hit@10: 98.5%
  • MRR: 0.540
  • Mean Turns to Conversion: 2.06
  • Technical Score: 0.8333

Compared with the baseline:

  • Hit@10 improved from 12.5% to 98.5%
  • MRR improved from 0.0680 to 0.5400
  • Mean turns decreased from 9.81 to 2.06
  • Technical Score improved from 0.1067 to 0.8333

The retained evaluation contains all 200 sessions, 409 turns, and 816 completed SDK calls. Failed turns were not removed from the results.

Beyond the score, we are proud that the system is inspectable. For any recommendation, we can reconstruct the state patch, active constraints, selected retrieval policy, route overlap, constraint evidence, ranking signals, and final dialogue decision.

What we learned

We learned that conversational shopping quality depends as much on state correctness as on model size.

A powerful reranker cannot recover a product that was lost because the system forgot an earlier constraint or misunderstood an intent override. Conversely, coarse retrieval should optimize for recall, while the final ranker should own precision.

We also learned that:

  • Missing evidence must not be treated as negative evidence.
  • LLM proposals should be validated before they change system behavior.
  • Dynamic retrieval involves both route weights and computation budgets.
  • User-profile preferences should remain soft and subordinate to explicit session intent.
  • Benchmark scores are more convincing when accompanied by reproducible traces and honest failure accounting.

What's next for Narrow

Our next step is to connect the existing over-generality signal to a true pre-retrieval cutoff, allowing Narrow to skip expensive retrieval and reranking when a request is too broad.

We also plan to add:

  • Provenance-aware decay for old soft slots
  • Cross-session profile distillation with explicit privacy controls
  • Durable profile and session persistence
  • More diverse LambdaMART training targets
  • Better calibration of recommend-versus-clarify decisions
  • Broader human-like evaluations for ambiguous, contradictory, and changing requests

Our long-term goal is to make Narrow a shopping copilot that knows not only which products to retrieve, but also when to search, when to remember, and when to ask.

Built With

Share this project:

Updates

Submission history