Inspiration
Shopping is rarely a single, perfectly written search query.
A customer may begin with:
“I need black running shoes below $100.”
They may then say:
“Those options are not quite right.”
A few turns later, they may completely change their mind:
“Actually, forget the shoes. I need a blue rain jacket instead.”
Traditional e-commerce search engines struggle with this type of conversation. Searching only the latest message loses important context, while combining every message can preserve outdated or contradictory preferences.
We were inspired by a simple question:
Can a shopping agent understand not only what a customer says, but also what it should remember, what it should forget, and when it should ask for more information?
Our goal was to build a conversational shopping copilot that follows the customer’s evolving intent and retrieves the exact product they are likely to purchase in as few turns as possible.
What it does
Our project is a deterministic, multi-turn conversational product-search agent built for the Amazon clothing, shoes, and jewellery catalog.
At every turn, it converts the conversation into an active shopping state containing information such as:
- product category;
- price range;
- colour;
- material;
- style;
- use case;
- exclusions;
- hard constraints;
- soft preferences.
It then determines whether the customer is Buying, Browsing, in a Boundary/no-preference situation, or expressing Uncertain intent.
Each route behaves differently.
For a high-intent Buying request, the agent protects exact categories, price limits, attributes, and other hard constraints. For Browsing, it allows broader semantic and use-case matching. Boundary sessions use only safe evidence without inventing preferences.
The agent also understands conversational changes. New explicit instructions override older preferences, negated attributes are removed, rejected products are excluded, and a genuine intent change resets conflicting state.
When a customer gives generic feedback such as “those are not quite right,” the agent does not search those words. Instead, it returns to the last informative request, keeps the valid constraints, excludes products already shown, and generates a fresh result set.
It can also ask clarification questions, but only when the candidate pool is genuinely ambiguous and the expected answer is likely to improve the result. Recommendations are still returned while clarification is requested, so the conversation never stops unnecessarily.
How we built it
We began with the official BM25 starter agent.
The baseline searched each new message independently. On the 200 public development sessions, it achieved a TechnicalScore of 0.106710, with a Hit Rate@10 of 0.125.
Rather than replacing the baseline immediately with a large model, we first examined where and why it failed.
The biggest early issue was conversational memory. If a customer said, “Those products are not quite right,” the baseline searched the rejection sentence and forgot the original request. We therefore built a structured state machine that separates the customer’s active intent from historical conversational evidence.
Our architecture contains several stages:
1. Structured session state
The state layer records active slots, constraints, exclusions, rejected products, question history, and the source of each piece of evidence.
Current-turn instructions have the highest priority. Conversation history is used only when it remains compatible, while profile information is treated as a soft signal rather than a hard constraint.
2. Intent routing
A deterministic router classifies each turn as Buying, Browsing, Boundary, or Uncertain.
This allows the agent to change retrieval behaviour according to the customer’s current goal rather than applying the same search policy to every conversation.
3. Candidate retrieval
We implemented lexical, dense, and hybrid retrieval components and evaluated them through controlled experiments.
One of our most important findings was that semantically similar products were not always the exact products used by the evaluator. Dense retrieval could return products that looked reasonable to a person but still receive no credit because the target product identifier was missing from the Top 10.
Our final configuration therefore gives greater importance to deterministic lexical and catalog evidence while keeping the architecture modular enough to support alternative retrieval policies.
4. Deterministic reranking
The reranker operates on a bounded candidate pool instead of searching the entire catalog again.
It combines signals such as:
- lexical retrieval rank;
- category compatibility;
- price compatibility;
- exact and partial catalog phrases;
- attribute coverage;
- rare identifying terms;
- use-case matches;
- popularity evidence;
- historical evidence;
- agreement between independent signals;
- contradiction and exclusion penalties.
Hard-constraint violations receive strong penalties or are removed entirely.
5. Utility-aware clarification
Clarification questions consume an additional interaction turn, so asking more questions is not always better.
Our clarification controller estimates whether a missing attribute would meaningfully separate the leading candidates and whether the customer is likely to provide a useful answer.
It avoids repeated questions, tracks pending and declined questions, stops questions after an intent change, and never asks when there is insufficient time remaining in the session.
6. Validation and safe fallbacks
Before returning a result, a central validator checks that recommendations contain valid catalog identifiers, removes duplicates, enforces the requested Top-K size, and applies a safe fallback if an optional component fails.
The final runtime is fully local and deterministic. It does not require a hosted LLM or external inference API.
Challenges we ran into
Semantic relevance was not the same as evaluation success
Our initial assumption was that adding dense retrieval would automatically improve the agent. It often produced products that were semantically appropriate, but the evaluator rewarded the exact purchased product.
This forced us to align every component with the actual objective rather than optimizing only for results that looked reasonable.
Remembering and forgetting were equally important
A useful shopping agent must preserve valid information across turns, but it must also remove preferences that the customer has replaced.
Intent Override sessions were especially difficult because retaining too much history created contradictions, while clearing all history removed potentially useful evidence.
We addressed this by separating active state from bounded historical evidence.
Generic feedback contained almost no searchable information
Messages such as “not quite right” communicate rejection but do not describe the desired product.
Recognizing these turns as feedback—and reusing the last informative request while excluding previous results—became one of our largest improvements.
Clarification had a measurable cost
A natural-sounding question could still reduce the final score by wasting a turn or asking for information that the customer could not provide.
We learned to treat clarification as an optimization decision, not a default chatbot behaviour.
Avoiding overfitting
The public development set contains only 200 sessions, while the organizer’s private evaluation uses separate users and target products.
We therefore froze promising configurations, tested one controlled change at a time, and rejected later variants when they did not pass our development gate.
Accomplishments that we are proud of
Our final public-development results were:
| Metric | BM25 baseline | Final agent |
|---|---|---|
| Hit Rate@10 | 0.125000 | 0.870000 |
| MRR | 0.068034 | 0.519609 |
| MTTC | 9.810 | 3.845 |
| Efficiency | 0.1190 | 0.7155 |
| TechnicalScore | 0.106710 | 0.733983 |
This represents approximately a 6.88× TechnicalScore improvement over the original BM25 baseline.
We are particularly proud that the improvement did not come from simply adding a larger model. It came from understanding the evaluation objective, building a more accurate representation of conversational state, and making every retrieval and clarification decision measurable.
The final system is deterministic, reproducible, fully in-memory, and requires no paid model API.
What we learned
The most important lesson was that conversational memory is not simply the ability to remember everything.
Good memory is controlled memory.
The agent must preserve evidence that is still useful, remove evidence that has become invalid, and distinguish an active customer requirement from a historical hint.
We also learned that system complexity must be earned through evidence. Dense retrieval, hybrid fusion, additional ranking features, and clarification policies were each evaluated independently. Components were kept only when they improved the measured result.
Finally, we learned that a shopping conversation is a sequence of ranking decisions. Even asking a question is part of ranking because it changes the evidence available and consumes one of the limited interaction turns.
What’s next
Our next priority is improving generalization to unseen users and products.
We would expand our held-out evaluation using catalog-stratified, non-public target samples so that ranking changes can be tested without repeatedly optimizing against the official public sessions.
We would also focus on Intent Override, which remains the most difficult scenario. Future improvements could include stronger category-transition detection, better separation of reusable historical evidence from outdated constraints, and more accurate calibration of clarification utility.
A production version could additionally provide explanations such as which attributes caused a product to be included or excluded. This would make the system easier to debug and give customers greater confidence in its recommendations.
Ultimately, our project shows that a capable shopping copilot does not only need to understand products.
It needs to understand how people change their minds.
Our agent knows what to remember, what to forget, and when to ask—helping customers reach the right product in fewer turns.
Built with
Python, Git, GitHub, Codex-assisted development, the official TechJam Agent interface and deterministic evaluator, and the frozen Amazon Reviews 2023-derived product catalog.
The final agent runs locally and does not require external model inference APIs.
Built With
- codex
- openai
- python
Log in or sign up for Devpost to join the conversation.