RatLab Shopping Copilot

RatLab Shopping Copilot is a fully local conversational product-search and recommendation agent developed for TikTok TechJam 2026 Statement 4: Shopping Copilot: AI Conversational Search and Recommendations. Traditional e-commerce search works well when users already know exactly what they want, but performs poorly when their needs are vague, evolve over several turns, or are expressed through natural conversation. Our solution is designed to handle the four main customer behaviours in the challenge: high-intent buying, exploratory browsing, users who change their mind midway, and users who remain vague throughout the conversation.

The system searches the organizer-provided frozen catalog of 50,000 Clothing, Shoes and Jewelry products derived from Amazon Reviews 2023, while maintaining conversational state and attempting to surface the customer's target product in as few turns as possible. The competition also provides 200 public development sessions for local evaluation, with a separate private evaluation set.

How our solution addresses the problem statement

Rather than relying on a single retrieval algorithm, RatLab Shopping Copilot combines several complementary retrieval and reasoning mechanisms. First, a multi-turn dialogue state tracker accumulates information such as product category, brand, color, material, size and budget. It supports negation, removal and explicit intent overrides, allowing a user to say something such as "actually, not red" or change their requirements without restarting the search. An intent-aware router then determines whether the conversation is currently closer to Buying or Browsing intent. This changes how the system combines lexical and semantic retrieval signals. For retrieval, we use a hybrid pipeline consisting of:

  • weighted SQLite FTS5 full-text search for fast, precise lexical retrieval;
  • optional local FastEmbed semantic embeddings using BAAI/bge-small-en-v1.5;
  • intent-dependent fusion of the sparse and dense results.

The resulting candidate pool is passed through a constraint-aware deterministic reranker, which considers active positive and negative preferences, budget constraints, recency of conversational information, catalog evidence and stable tie-breaking.

We additionally built a global catalog-evidence index over all 50,000 products. It indexes user-expressible facts including titles, taxonomy segments, feature descriptions, structured product details, materials, colors, stores and price information. This allows an exact clue mentioned by a customer to recover the correct product even if it did not appear in the initial retrieval candidate pool. Our final agent prioritizes the customer's original wording over potentially noisy automatically extracted slots.

Finally, the agent uses proactive clarification and progressive recommendation breadth. During the first few turns it asks structured questions and only exposes its strongest candidate while information is still accumulating. Once sufficient information has been gathered, it expands its recommendations to the full Top 10. This helps reduce unnecessary conversation and directly targets the challenge's MTTC efficiency metric.

An important design decision was to keep the submitted agent fully local and deterministic. It does not call an external LLM or paid inference API during evaluation, does not require an external vector database, and reports zero prompt and completion tokens.

Development tools used

Our development workflow used: Python 3.12 for implementation and experimentation

Git and GitHub for version control, issues, branches, pull requests and team collaboration

GitHub Desktop for local repository and branch management

Python virtual environments for reproducible dependency management

The organizer-provided deterministic local evaluator for continuous evaluation and regression testing

Python's unittest framework, with 87 unit and integration tests in the final submission covering dialogue state, retrieval, routing, reranking, clarification, overrides and catalog evidence.

APIs used

No external model or commercial APIs are required at evaluation time.

The final agent is deliberately offline after its local resources have been installed and cached. It does not use OpenAI, Google, hosted embedding services or external vector-database APIs.

The implementation conforms to the official TikTok TechJam Agent API contract, exposing the required reset() and respond() interface to the organizer's evaluator.

Libraries and frameworks used

The main libraries and technologies used are:

FastEmbed 0.8.0 — local semantic embedding inference

BAAI/bge-small-en-v1.5 — 384-dimensional text embedding model used for the dense retrieval track

NumPy 2.2.6 — vector storage, normalization and similarity computation

SQLite FTS5 — fast in-memory/local full-text lexical retrieval

ONNX Runtime, through FastEmbed — efficient local embedding inference

Python standard-library components — parsing, state management, indexing and evaluation.

We intentionally avoided heavyweight external vector databases or remotely hosted inference infrastructure so that the complete retrieval pipeline can operate locally.

Datasets and assets used

We used only competition-authorized and publicly available resources: The organizer-provided frozen 50,000-product Amazon Reviews 2023 Clothing_Shoes_and_Jewelry catalog

The organizer-provided 200 labelled public development conversations used for local testing and optimization

The official TikTok TechJam local evaluator and starter repository

Locally generated embeddings of the competition catalog using BAAI/bge-small-en-v1.5

No private evaluation sessions, private organizer labels or mock products were used by the agent.

The source catalog is derived from Amazon Reviews 2023, published by the McAuley Lab at UCSD, with parent_asin used as the product identifier. The competition data contains text and structured product metadata rather than multimodal content.

Results

On all 200 released public development sessions, our final configuration achieved:

Metric Result
Hit Rate@10 1.000000
MRR 0.974048
MTTC 2.070000 turns
Efficiency 0.893000
TechnicalScore 0.970814
Model tokens 0

It achieved a 100% Hit Rate@10 across Buying, Browsing, Intent Override and Boundary scenarios on the public development set. These results are public-set measurements and are not presented as a guarantee of performance on the organizer's unseen private evaluation set.

Closing

RatLab Shopping Copilot demonstrates that a conversational shopping agent does not necessarily require a large hosted LLM to achieve strong retrieval performance. By combining dialogue-state tracking, intent-aware sparse/dense retrieval, exact catalog evidence, deterministic reranking and proactive clarification, we built a lightweight system capable of adapting to evolving customer requirements while remaining fast, reproducible and completely local.

Built With

Share this project:

Updates

Submission history