Inspiration

The starter agent we were given scores 0.10671. It searches by ORing every word in the customer's message together, so a request for "black cotton running shirt" matches anything containing the word black. We wanted to know how much of the gap between that and a good agent was about better search, and how much was about better conversation.

The answer turned out to be almost entirely the second one, and finding that out changed how we spent the rest of the four days.

What it does

Kwekers Shopping Copilot finds one hidden product in a 50,000-item clothing catalog by talking to a customer, in ten turns or fewer.

Each turn it returns ten ranked product IDs and asks one clarifying question. It accumulates what the customer has told it across turns, detects when they change their mind, and reorders its candidate pool as evidence arrives. The conversation ends when the target appears in its top ten.

On the 200-session public development set it finds the target in every session, in 2.06 turns on average, scoring 0.891084 against the organiser's 0.10671 starter baseline.

One number needs a caveat we state everywhere it appears. HitRate@10 of 1.000 is cumulative across up to ten turns with ten products per turn. It is not single-turn accuracy. MRR of 0.708 is the figure that describes ranking quality, and it means the target usually sits around rank 1.4 when found.

The scored path calls no language model. It runs offline in roughly 30 seconds with zero token cost.

How we built it

The scored path is one deterministic pipeline.

Regex-first parsing extracts constraints from the simulator's message templates, and a scenario router classifies each conversation as buying, browsing, intent override, or boundary from the same signals.

Accumulated dialog state keeps constraints across turns rather than replacing them. When the customer overrides an earlier preference, we demote it to a soft signal instead of deleting it, because the target product does not change when someone re-prioritises, so the old metadata still helps rank it.

BM25 retrieval is the only route that can add a candidate. Field-weighted SQLite FTS5 with a three-tier strictness cascade, exact phrase then AND then OR, so a stricter match always outranks a looser one. It returns a pool of 500.

Evidence layers reorder that pool but never extend it. An exact-match index intersects catalog-verbatim constraint strings and adds a flat boost of 0.35. A category-bucket route adds 0.10. A dense bi-encoder route is fully wired but ships disabled at weight 0.0, because we measured that it cost 0.075 end to end.

Freshness excludes already-shown products so the agent does not repeat itself for ten turns, and clears that exclusion on an override so a target hidden by rotation can resurface.

An optional language model layer, off by default, handles unfamiliar phrasing and writes the customer-facing explanations. It never touches product selection. Any constraint the model proposes must pass our normalisation and match an entry in the catalog index before acceptance, so it can only select strings that already exist and cannot invent one. The client accepts only a free-tier model, times out after 3 seconds, and returns nothing on any error so the caller always has a deterministic fallback.

With every flag off the agent makes zero network calls, which a test asserts alongside the score.

We measured everything against the real evaluator rather than proxy metrics, on a scenario-stratified 140/60 split, with a paired bootstrap for significance. At n=200 the unpaired standard error is 0.0096, so differences under roughly 0.019 are indistinguishable. We adopted a rule of preferring the simpler configuration unless a more complex one clears that band.

Challenges we ran into

A SQLite non-determinism we found ourselves. Our BM25 query had no secondary sort key, so exactly-tied scores could resolve to a different row order across SQLite builds. On the public set this changed one session from rank 7 to rank 8 on the same hit turn, moving our frozen score from 0.891111 to 0.891084. We fixed it and confirmed the exact blast radius before shipping, since it moved a number we had already documented as frozen. A slightly lower score that reproduces on any machine is better evidence than a higher one that depends on runtime accident.

An exact-match veto bug silently discarding correct products. Our exact-evidence route treated any constraint it could not index, colors in particular, as a hard veto on the entire candidate set rather than as no evidence. Moving to three-state semantics, where unsupported constraints skip, supported zero-match constraints veto, and supported matches intersect, moved the score from 0.877011 to 0.891111. It was the single largest correctness fix of the project.

A proxy metric that pointed the wrong way. Dense rescoring looked good in isolation, worth +0.078 on turn-1 Recall@10, but cost 0.075 once it ran inside the full ten-turn conversation. By turn 2 or 3 the accumulated constraints already order BM25 well, and rescoring shuffled a good pool against a noisier signal. We swept turn gates to see whether applying it only early would help. No gate beat not using it at all.

A normalisation function stripping the characters its own patterns needed. Our text cleaner removed colons and semicolons before the marker patterns looking for phrases like "what matters is" ever ran, so multi-constraint replies never split correctly for any caller feeding it raw templated text.

A contested decision we reported honestly instead of resolving in our favour. The category-bucket term is significant against a dense-enabled baseline, +0.0122 with CI [0.0065, 0.0182], but not against the configuration we actually ship, +0.0069 with CI [-0.0039, 0.0185]. Its value appears to depend on whether dense retrieval is also active. We report both framings.

Accomplishments that we're proud of

Seven ideas measured and rejected. Dense retrieval as a route, dense rescoring, character n-grams as a reranker, near-duplicate suppression, category filtering, BM25 field-weight tuning, and a rotating question policy. Each was plausible, each was tested, and each was dropped with its measured delta recorded. The final system is simple because those measurements say it should be.

We built an adversarial harness to attack our own system. Four levels of synthetic paraphrasing that rewrite customer messages while preserving meaning. The shipped configuration scores 0.891111 at level 0 and 0.419598 at level 4, where constraints are genuinely reworded. At level 4, intent-override sessions collapse from 1.000 to 0.033. That is a large drop and we report it, because the harness exists so we would know how far the score falls before it happens rather than after.

A frozen baseline test that asserts both the score and a zero LLM call count. It is what makes the claim that our optional model layer is genuinely optional something we can prove rather than assert.

A paired bootstrap that overturned our own heuristic. Our unpaired noise band called a +0.0282 delta noise. Because configurations run on the same sessions and are correlated, the paired test found it significant with 95% CI [0.0078, 0.0483]. We would have made a worse decision using the simpler test.

What we learned

This is a ranking problem, not a retrieval problem. Sweeping BM25's pool depth on turn 1 gave Recall@10 of 0.221, Recall@100 of 0.643, and Recall@500 of 0.950. Almost everything the agent needs is already inside a pool of 500 candidates on the first turn. The loss happens between rank 10 and rank 500. That single measurement redirected the whole project away from adding retrieval routes.

Conversation handling moved the score more than any retrieval route did. Excluding already-shown products and clearing that exclusion on an override raised the score by 0.044 and HitRate by 0.055. On override sessions specifically, demoting a superseded preference instead of deleting it was the difference between 0.333 and 1.000 HitRate. Meanwhile adding a second retrieval route made things worse, with BM25 plus character n-grams scoring 0.585 against BM25 alone at 0.630.

Isolated metrics lie about full-system behaviour. The dense rescoring result is the clearest case. It improved the number we were measuring and hurt the number that mattered. We stopped trusting any proxy without also running the full evaluator.

Getting better made us more fragile. The gain from 0.855 to 0.891 came almost entirely from exact-match evidence, which is the component that most depends on the simulator quoting catalog metadata close to verbatim. Our robustness under paraphrasing got worse as our score got better. That is a real trade-off and we would rather state it than have a judge find it.

What's next for Kwekers Shopping Copilot

Robust constraint extraction. The exact-match route is worth 0.155 MRR and needs verbatim catalog text. On real customer traffic, where people say "soft, not see-through" instead of "95% cotton, 5% spandex," it would rarely fire. The optional language model layer is our first step at this, and the adversarial harness is how we would measure whether it is working.

Abstention. The agent always returns ten recommendations and never says it is unsure. We built a read-only confidence layer as groundwork and measured that it sits flat at around 0.008 across all turns. We diagnosed why. The pool saturates at 500 candidates regardless of disclosure, only about one candidate per turn earns the exact-match boost, and two-thirds of the pool shares the same flat category boost, so the distribution stays high-entropy. The fix is to normalise scores within a truncated pool before computing entropy rather than after.

A rejection channel. The simulator can reveal more preferences or state it has none, but it cannot say "not that one," which is arguably the most useful signal in real shopping. Our agent has no way to receive it even if it existed.

Recommendation behaviour that suits a person rather than a metric. Our candidate rotation is optimal for HitRate@10 but deliberately withholds a good match so it can count later. A production version would persist high-confidence items and rotate only the uncertain tail.

How our solution tackles the problem

Our project is an intelligent shopping agent that understands whether a user is buying with specific requirements or casually browsing. It combines keyword, category, exact-constraint and optional semantic retrieval to recommend products from the Amazon Reviews 2023 dataset. Across multiple turns, it remembers preferences, updates constraints, handles sudden intent changes and asks clarifying questions when requests are too broad. The system is optimized for accurate recommendations and fewer interactions, measured using Hit Rate@K, MRR and Mean Turns to Conversion.

Built With

Share this project:

Updates

Submission history