Inspiration
“I don’t want blue” and “actually, make it red instead of blue” both remove blue from the current request—but only one means the shopper dislikes blue. That distinction inspired IntentCart. Finding the right product often takes a conversation. Shoppers begin with incomplete ideas, answer follow-up questions, reject options, and revise earlier preferences. A useful shopping agent must understand which requirements remain active, which were explicitly rejected, and which were simply replaced.
We wanted to discover how far structured conversation memory, complementary retrieval strategies, and explainable ranking could go without calling a generative model on every turn.
What it does
IntentCart receives an anonymized preference profile and one customer message at a time. On every turn, it can ask one structured clarification question, return ranked product recommendations, or do both.
The agent can:
- Remember requirements across multiple turns.
- Interpret short answers using the previous question.
- Distinguish hard requirements from soft preferences.
- Understand explicit negatives such as “avoid leather.”
- Recognize answers such as “any color is fine” without adding false constraints.
- Replace outdated preferences when the shopper changes their mind.
- Avoid repeating products under an unchanged intent.
- Return valid, unique catalog identifiers.
IntentCart keeps active, negative, declined, and superseded preferences separate.
For example:
- “I don’t want blue.” Blue becomes an explicit negative.
- “I don’t mind the color.” Color is removed as a requirement.
- “Actually, make it red instead.” Blue becomes superseded history, while red becomes active intent.
This prevents outdated preferences from silently affecting later recommendations.
How we built it
Customer message + anonymized profile
|
v
Dialogue-state manager
(preferences, negatives, overrides)
|
+--> Broad BM25 retrieval --------+
+--> Strict AND retrieval --------+
+--> MiniLM semantic retrieval ----+
(optional) |
v
Candidate union + deduplication
|
v
Constraint-aware ranking
|
v
Clarification + ranked recommendations
Conversational intent memory
starter/dialog.py maintains isolated state for each customer session.
Instead of concatenating the entire conversation into increasingly noisy text, it stores active requirements, explicit negatives, superseded values, declined attributes, pending questions, and conversation history separately.
The pending question helps interpret short answers. If the agent just asked about color, “blue” becomes a color constraint. If the customer instead says “actually, forget that—I need something waterproof,” the override cancels the pending color question.
Interpretation follows a deliberate order: declines first, intent overrides second, and ordinary preference information last. The resulting structured decision contains the current query, category, constraints, priorities, negative values, override status, and next clarification attribute.
Multi-route lexical retrieval
starter/retrieval.py loads the catalog into an in-memory SQLite FTS5 index.
The retriever searches four independent views of the current intent:
- Consolidated search query
- Active constraints
- Product category
- Low-weight customer-profile evidence
BM25 scores are normalized and combined through weighted fusion. Products found through multiple routes receive corroborating evidence. A separate strict route requires every disclosed searchable concept. Broad retrieval protects recall, while strict retrieval contributes complete, high-precision matches. The agent unions both pools before ranking, so an empty strict result cannot remove useful broad candidates.
Optional semantic retrieval
starter/embedding_retrieval.py optionally uses sentence-transformers/all-MiniLM-L6-v2.
This route helps when the shopper and catalog express the same need differently—for example, “shoes for standing all day” and “long-shift comfort with all-day cushioning.”
Product embeddings are computed once and stored in a reusable, catalog-fingerprinted cache. Only the current intent is encoded during each new search.
Semantic candidates are added to the lexical pool rather than replacing it. This allows MiniLM to expand recall while exact category, color, material, size, brand, and budget evidence remains in control.
If the model or cache is unavailable, IntentCart continues through its complete lexical pipeline.
Constraint-aware ranking
starter/ranking.py reranks the combined candidate pool using:
- Retrieval strength
- Current-intent agreement
- Active-constraint agreement
- Retrieval-route evidence
- Profile overlap
- Product-quality evidence
- Budget, size, and exclusion signals
Current intent receives more influence than historical profile evidence.
The ranker also distinguishes missing metadata from a confirmed mismatch. A product with no known price is not treated the same as a product known to exceed the customer’s budget.
Finally, starter/agent.py selects unseen recommendations, chooses the next clarification field, and returns the exact required Agent API response.
What makes IntentCart different
IntentCart’s main contribution is not simply adding embeddings. It is the coordination of explainable components that solve different failure modes:
- State-aware intent replacement: Replaced values are removed without automatically becoming dislikes.
- Complementary retrieval: Broad, strict, and semantic routes contribute candidates independently.
- Lexical-first semantics: Semantic similarity expands vocabulary coverage without weakening exact constraints.
- Unknown-aware ranking: Missing catalog metadata remains eligible instead of becoming a false mismatch.
- Deterministic behavior: Questions, ranking, fallbacks, and tie-breaking are reproducible.
The system does not call a generative LLM or hosted inference API during evaluation. It requires no API credentials and reports zero prompt and completion tokens.
Challenges we ran into
Intent is state, not accumulated text
Our early approach treated the conversation as one growing query. It failed when customers changed their minds or said that an attribute did not matter. Separating active, negative, declined, and superseded values made intent transitions explicit and testable.
Balancing recall and precision
Broad retrieval found more possible targets but admitted partial matches. Strict retrieval improved precision but could return nothing when metadata was incomplete. We retained both approaches and unioned their candidates before ranking.
Making semantics useful without weakening constraints
At a conservative calibration, MiniLM improved one public session by moving the correct target from rank eight to rank six. Increasing semantic influence caused more regressions than improvements. This taught us that embeddings should complement lexical and structured evidence rather than dominate them.
Optimizing competing metrics
Returning ten products immediately can find targets sooner, improving MTTC. Smaller early recommendation batches allow later answers to improve the target’s rank, helping MRR. We evaluated these policies separately and selected the tradeoff that performed best under the competition’s Technical Score formula.
Results
On the organizer-provided 200-session public development set:
| Configuration | Hit Rate@10 | MRR | MTTC | Technical Score |
|---|---|---|---|---|
| Weak starter | 0.125000 | 0.068034 | 9.810 | 0.106710 |
| Lexical/strict IntentCart | 1.000000 | 0.861851 | 2.765 | 0.923255 |
| Hybrid, semantic scale 0.75 | 1.000000 | 0.862060 | 2.765 | 0.923318 |
Most of the improvement came from dialogue state, multi-route lexical retrieval, and constraint-aware ranking. MiniLM produced a smaller, measured improvement, which we report as evidence of complementary semantic recall rather than a guaranteed hidden-set advantage.
Additional robustness testing
We also evaluated the lexical configuration on a locally generated 500-session synthetic stress test with unique targets and zero target overlap with the public set:
| Metric | Result |
|---|---|
| Hit Rate@10 | 0.982000 |
| MRR | 0.596803 |
| MTTC | 2.930000 |
| Technical Score | 0.831441 |
The synthetic set is supplementary validation. It is not organizer-provided and is not an official private-evaluation result. Public results are also not a guarantee of hidden-set performance.
Accomplishments we are proud of
- Increased public Technical Score from 0.106710 to over 0.923.
- Reached Hit Rate@10 1.0 across all 200 public sessions.
- Explicitly handled Buying, Browsing, Boundary, and Intent Override scenarios.
- Built broad, strict, and optional semantic retrieval routes.
- Added 151 deterministic unit and integration tests.
- Created a reproducible 500-session stress-test generator.
- Kept external inference calls, credentials, API tokens, and API cost at zero.
- Preserved a complete lexical fallback when semantic retrieval is unavailable.
What we learned
We learned that correct conversational memory can matter more than adding a larger model. We also learned that:
- Replacing a preference is different from rejecting it.
- Pending-question context is essential for interpreting short answers.
- Broad and strict retrieval solve complementary problems.
- More semantic influence is not automatically better.
- Missing metadata must be separated from confirmed mismatches.
- Deterministic systems are easier to test, reproduce, and explain.
- Public, synthetic, and hidden evaluation results must be reported separately.
What’s next for IntentCart
With more time, we would:
- Improve typo tolerance and unusual paraphrase handling.
- Support multilingual shopping conversations.
- Extend candidate-aware questions to more product attributes.
- Add explanations showing which constraints each recommendation satisfies.
- Evaluate a lightweight cross-encoder on the final candidate pool.
- Build an interactive interface for user testing.
Development tools
We used Visual Studio Code, macOS Terminal, Python command-line tools, Git, and GitHub. We did not require Google Colab or Jupyter Notebook.
APIs used
No external inference API is used during evaluation. The optional MiniLM model is downloaded from the Hugging Face Hub during setup. After provisioning, it runs locally without an inference API, API credentials, or billable tokens.
Datasets and assets
- Amazon Reviews 2023 by McAuley Lab, UCSD
- Clothing_Shoes_and_Jewelry category
- Frozen catalog of 50,000 products
- Organizer-provided 200-session public development set
- Locally generated 500-session synthetic stress-test set
- sentence-transformers/all-MiniLM-L6-v2
Log in or sign up for Devpost to join the conversation.