Inspiration

Online shopping filters are rigid, but real customers aren't. Nobody types material=leather AND closure=buckle. They say "something like my old belt, but nicer," change their mind mid-conversation, and sometimes don't know what they want until they see it.

That gap is what drew us in. We wanted to build a shopping assistant that treats a conversation as evidence to reason over, not a string to match: LLMs and embeddings to understand shifting intent and ask useful follow-up questions, and fast exact retrieval to do the heavy lifting underneath. The goal was an agent that adapts to changing preferences, handles uncertainty, and still finds the exact right product in a catalog of 50,000.

What it does

Our agent talks with a shopper and tries to find their hidden target in a 50,000-item Amazon catalog.

On the official evaluator, it scored 0.9755 TechnicalScore. It found the target in all 200 public sessions and placed it at rank #1 in almost every case, with 0.983 MRR and an average of 1.98 turns to conversion. The starter baseline scored 0.1067.

That 0.9755 result comes from the offline mode and reproduces exactly on every run. Our measured run with the live LLM lanes enabled scored 0.9709.

On each turn, the agent:

  • extracts what the shopper just revealed
  • updates its memory of their preferences
  • decides whether they are buying or browsing
  • retrieves and ranks candidates through five parallel lanes
  • decides whether to show results or ask one more question

That final decision matters more than we expected. A premature answer can lock in a poor rank, while one well-chosen question can isolate the target.

The agent tracks material, color, brand, category, size, and budget separately. If the shopper changes the color, only the color changes. Everything else stays intact. It calls the LLM only when cheaper methods cannot settle the result, which keeps most turns fast, traceable, and nearly free.

How we built it

Why blindly trust an expensive, hallucination-prone AI call when hard math can find the exact answer in milliseconds? That philosophy drove the architecture.

1. Attribute-scoped state management. The memory tracker keeps each preference in its own slot. If a shopper says, "Actually, show me blue Nike shoes instead of red," the agent changes the color to blue while preserving the brand, size, category, and budget. Forgetting the entire conversation after an override sounds safe, but it throws away evidence that is still valid. We tested that approach. It made the score worse.

2. The 5-lane hybrid retrieval engine. Five retrieval lanes run over the catalog at once:

Exact-constraint lane — stated requirements are treated as literal evidence and used to eliminate every product that couldn't have produced them (the highest-precision signal we have)

Category + keyword lanes (BM25) — instant lexical matching over titles, features, details Prior lane — review-count popularity, so plausible products surface when evidence ties Dense semantic lane — embeddings that understand open-ended vibes like "something cozy for a winter flight", for when wording and catalog text don't overlap

The lanes are fused in strict tiers: categorical evidence outranks similarity scores, so a high cosine score can never outvote a literal constraint match.

3. Four-stage execution gating + Expected Information Gain. A powerful engine also needs to know when to hit the brakes, so four gates block unnecessary work:

  • Gate 1 skips retrieval when a message carries no new information.
  • Gate 2 shuts off the expensive semantic lane when hard evidence has already isolated a single candidate. No similarity score can reorder a one-element list.
  • Gate 3 - the over-generality gate. If the request is too broad ("I want shoes"), showing products wastes a turn. We pause, compute Expected Information Gain over the remaining pool, and ask the one question that cuts it down fastest.
  • Gate 4 - the emission gate. The evaluator freezes your rank the moment the target appears in your top-10, so revealing a list early can lock in a bad rank forever. We derived the break-even from the scoring weights: rank #1 is worth roughly 7× more than one saved turn. While uncertain, the agent shows a single best guess (which can only ever hit at rank #1) and keeps asking; it opens up to the full ten only when the evidence is exhausted.

4. Gated LLM reranker + demo. The system spends an LLM call only when the final candidates remain genuinely tied. A reranker through OpenRouter then reorders the shortlist and tries to put the strongest match first.

The pipeline supports two modes.

The offline mode is pure Python, deterministic, and costs nothing to run. The live mode adds LLM extraction, dense vector retrieval, and final reranking through OpenRouter. Our measured live run scored 0.9709.

So why keep the AI lanes if the deterministic mode scores slightly higher on the evaluator?

Because the offline system knows the evaluator's language unusually well. Real shoppers do not repeat catalog fields word for word. In our paraphrase stress test, the offline score fell by 40 percent as the simulator began speaking more naturally. The semantic and LLM-backed lanes recovered much of that loss.

5. Live Demo (Extra) A bundled web demo replays real evaluator sessions with the agent's internals visible: the candidate pool collapsing, constraints accumulating, and the emission gate deciding when to commit.

Challenges we ran into

The first challenge was translating natural language into the evaluator's attribute vocabulary: category, material, color, size, style, brand, budget, and use case.

Our team came from different technical backgrounds, and we began with different ideas about how extraction should work. Reaching one shared model took several whiteboard sessions and more trial and error than we expected.

The harder lesson was that when the agent answers matters as much as what it retrieves. Our early retrieval results looked strong, but the final score did not. The agent was exposing top-ten lists too soon and freezing the target at mediocre ranks.

The turning point came when we stopped treating emission as a presentation choice and modeled it as a scoring decision. We derived the threshold from the evaluator's weights instead of guessing.

Live model calls added another set of problems: non-deterministic outputs, latency spikes, and the chance of a broken credential on judging day. We designed every model-backed stage to be optional. If a request times out or a key stops working, the deterministic core takes over.

Accomplishments that we're proud of

  • TechnicalScore 0.9718 vs a 0.1067 baseline. 100% hit rate across all four scenario types (buying, browsing, intent override, boundary), MRR 0.9746, MTTC 2.03 turns to conversion; 0.9709 measured with the live LLM lanes enabled.

  • A reproducible spine. The offline core scores exactly 0.9718 on every run — judges can verify it to the digit with one command. The LLM lanes sit on top as genuine capabilities: gated to fire only where they add value, with graceful degradation when they can't.

  • A measured ablation for every claim. Every component in this write-up has a number behind it, produced by the official evaluator — including what happens when we remove it.

  • Nearly free to run. Sub-millisecond turns in the core, NumPy as the only dependency, fractions of a cent per session when the LLM lanes are on, $0 when they're off.

  • Robustness we tested, not assumed. We built a paraphrase stress harness that rewords the simulated customer at increasing severity; the LLM extraction and dense lanes recover a large share of the score that exact matching loses when the customer stops quoting catalog wording.

We deliberately balanced leaderboard performance against reliability and extensibility — this architecture is one an actual e-commerce team could take toward production.

What we learned

A good solution starts with understanding the problem before writing code. We learned to turn the way people naturally describe preferences into structured product attributes; to combine LLMs, embeddings, exact matching, and adaptive ranking under real cost and latency constraints; and to let measurement, not intuition, settle design arguments (several "obviously good" ideas, like a blanket intent router, measurably hurt and were cut). We also learned to treat non-determinism as a design constraint: anything a live model touches needs a deterministic fallback and a number proving the fallback holds up. Most importantly, we learned how a team from different technical backgrounds turns an unfamiliar problem into a working, measurable system.

What's next for TIKTOK JONG JAROEN

The clear next step is promoting the paraphrase-robust path from fallback to first-class: real shoppers never quote catalog text, so LLM-based constraint extraction with semantic matching becomes the primary parser, with the exact-evidence engine and the emission policy unchanged underneath. Beyond that: testing against real humans rather than a simulator, approximate-nearest-neighbor indexes to scale past 50,000 products at the same latency, image understanding so a shopper can just send a photo of the belt they're replacing, and shipping the demo as an actual storefront widget.

Built With

+ 11 more
Share this project:

Updates

Submission history