Inspiration
Devpost Project Submission: Shopping Copilot
Project Title: Shopping Copilot: High-Precision Conversational E-Commerce Search Engine
Tagline: Turning 50,000-product conversational search into a 2.4-turn game of 20 Questions with zero latency and zero API cost.
Track: Track 4 (Conversational E-Commerce Search Challenge)
GitHub Repository: https://github.com/Bread04/TIKTOKJAM-
Demo Video (YouTube): https://www.youtube.com/watch?v=PVj8fxMgiro
1. Project Description and Problem Statement
The Problem
On live shopping and social video platforms like TikTok Shop, users rarely search with boolean keyword queries. Instead, they discover products through natural, multi-turn conversational dialogue:
- "I need an outfit for an outdoor summer wedding"
- "Something breathable and under $80"
- "Actually, ignore the red one, let's do navy blue instead"
Standard keyword search engines and naive baseline agents fail in conversational settings:
- Stateless Search: Treating each turn as an isolated query throws away context and constraints gathered earlier in the conversation.
- No Clarification Policy: Without asking targeted clarifying questions, an agent guesses blindly across a 50,000-product catalog.
- Repeated Recommendations: Re-showing previously rejected items wastes top-10 slots.
- High Failure Rate: The baseline starter agent failed 87.5% of sessions (Hit Rate@10 of 12.5%) and required 9.81 turns on average.
How Our Solution Addresses the Problem
Shopping Copilot treats conversational search as a game of 20 Questions over a 50,000-product catalog.
Rather than sending raw chat history to a slow, costly external LLM, Shopping Copilot runs a stateful, in-memory pool-refinement pipeline:
- Multi-Turn State Accumulation (
_SessionState): Tracks extracted constraints (category, color, material, size, budget, style), user profile tags, and conversation terms with recency weighting across all turns. - Hard Facet Filtering: Applies recognized product attributes (such as material, color, and price bounds) as strict candidate exclusions before ranking, preventing cross-category mismatches.
- Information-Gain Clarification (
ask_attribute): Evaluates surviving candidates and picks the attribute question that splits candidate distributions most evenly across attribute dimensions. - Guess-Who Elimination and Title Deduplication: Confirmed misses (
shown_asins) are removed from subsequent turns. Near-duplicate listings (>85% title text similarity) are pruned so top-10 slots test distinct product hypotheses. - Intent Override Recovery: Detects natural language preference shifts mid-dialogue, purges conflicting modifier slots, and resets candidate elimination bounds to avoid locking out the new target.
Results: Shopping Copilot improves Hit Rate@10 from 12.5% to 98.0%, raises MRR from 0.068 to 0.667, and finds target products in 2.41 turns on average (down from 9.81 turns), with $0.00 API cost and under 45 ms turn latency.
2. Before vs. After: Concrete Dialogue Example
Here is how the baseline agent and Shopping Copilot handle a multi-turn conversation from the benchmark:
User Prompt (Turn 1): "I'm looking for athletic running shoes. Black, size 10, under $60."
| Step | Baseline Agent | Shopping Copilot |
|---|---|---|
| Turn 1 Action | Returns 10 generic keyword matches. Does not ask any clarifying question. | Extracts category=shoes, color=black, size=10, budget=60. Queries FTS5 BM25 index, applies hard facet filters, and asks ask_attribute="material" to split the candidate pool. Returns top 10 deduplicated running shoes under $60. |
| Turn 2 Action | User: "Breathable mesh preferred." Baseline queries only "Breathable mesh preferred", loses shoe context, and shows mesh shirts. |
User: "Breathable mesh preferred." Updates BM25 candidate pool with mesh weighting. Hard filters remove non-mesh items. Excludes Turn 1 ASINs. |
| Final Result | Failed (0 hits in 10 turns) | Success on Turn 2 (Rank 1 Hit, Reciprocal Rank = 1.0) |
3. Development Tools Used
- Visual Studio Code: Primary code editor with Python and Pylance extensions for static type analysis and code navigation.
- Python 3.12 (64-bit): Execution environment utilizing built-in performance improvements and type annotations.
- Windows Terminal and PowerShell 7: Local shell environment for running benchmark test suites and interactive evaluation sessions.
- Git and GitHub: Version control for tracking code changes, test suites, and documentation across 24 ablation experiment runs.
4. APIs Used
Pure Offline Architecture ($0.00 API Cost, 0 External Dependencies)
Shopping Copilot uses zero paid external APIs (no OpenAI GPT-4o, Anthropic Claude, or Google Gemini).
Reasons for this Design Choice:
- Zero Operating Cost ($0.00 per session): External LLM calls cost \$0.01 to \$0.05 per conversation. At TikTok Shop scale (millions of daily conversations), this creates significant operational overhead. Our architecture costs \$0.00.
- Low Latency (<45 ms per turn): Cloud LLMs add 500 to 2,000 ms of network and generation latency per turn. Shopping Copilot executes each turn in under 45 milliseconds, completing the full 200-session benchmark in 84 seconds.
- Data Privacy: User chats, profiles, and catalog records remain entirely local in memory.
- Availability: Eliminates risks of rate limiting, token throttling, or external API outages during high-traffic shopping periods.
Internal Interfaces:
- Agent Lifecycle Interface (
starter/agent.py): ImplementsAgent(catalog_path),reset(session_id, user_profile), andrespond(session_id, user_message, turn, top_k) -> AgentResponse(recommended_asins, ask_attribute). - SQLite C-Interface (
sqlite3FTS5): High-speed in-memory full-text search API. - Evaluator Harness (
evaluator/local_evaluator.py): Multi-scenario benchmark runner simulating user buying personas.
5. Libraries and Frameworks Used
To ensure maximum reproducibility and zero dependency management issues, Shopping Copilot relies entirely on the Python Standard Library:
sqlite3(Full-Text Search): In-memory virtual table using SQLite FTS5 with BM25 ranking across 6 metadata fields (title,categories,features,details,store,description).re(Regular Expressions): Fast pattern matching for slot extraction (colors, materials, sizes, budgets, categories, and intent override signals).difflib(SequenceMatcher): Pairwise string similarity checking to identify and suppress near-duplicate catalog titles (>0.85 similarity threshold).collections(defaultdict,Counter): Multi-turn term frequency counting and attribute distribution histograms.math(math.log2): Shannon entropy calculations for selecting the most informativeask_attribute.jsonandgzip: Streaming decompression and parsing of the 50,000-product catalog and benchmark test files.unittest: Automated unit, boundary, and adversarial test suites (tests/test_agent_adversarial.py,tests/test_evaluator.py).
6. Datasets and Assets Used
1. Amazon Reviews 2023 (Clothing_Shoes_and_Jewelry)
- Catalog Size: 50,000 frozen real-world e-commerce product records (
data/catalog.jsonl.gz). - Fields:
parent_asin,title,price,categories,features,details,store,description.
2. Multi-Turn Public Benchmark Dataset
- Session Volume: 200 labeled multi-turn evaluation sessions (
data/public_set.jsonl). - Scenarios:
- Buying (80 sessions / 40%): Specific purchase intent with concrete constraints.
- Browsing (80 sessions / 40%): Broad exploratory discovery with open-ended initial queries.
- Intent Override (30 sessions / 15%): Preference reversals mid-conversation (e.g. changing color from green to black at turn 3).
- Boundary (10 sessions / 5%): Edge cases, unanswerable queries, and "no preference" responses.
7. System Architecture
+---------------------------------------------+
| Startup: In-Memory SQLite FTS5 |
| (50,000 Products, 6 Indexed Fields, BM25) |
+---------------------------------------------+
|
Per Turn (User Message) ------------------------------+
| |
v v
[1. Slot & Override Extraction] [2. Term Accumulation & BM25 Query]
- Extract: Category, Color, - Union accumulated conversation terms
Material, Size, Budget with recency weights + user profile
- Detect Intent Override - Execute BM25 retrieval -> Pool of 300
| |
+--------------------------------------+
v
[3. Hard Facet Filtering]
- Strict candidate exclusion: category mismatch, price tolerance (+/-30%),
and conflicting color/material/size slots
v
[4. Recency-Weighted Scoring & Verification]
- Score = BM25 + Recency Bonus + N-Gram Phrase Bonus + Category Match Bonus
- Exclude previously shown ASINs (Guess-Who Elimination)
v
[5. Near-Duplicate Title Deduplication]
- SequenceMatcher string distance (>85% similarity filter)
v
[6. Entropy Clarification & Final Output]
- Pick ask_attribute with highest information gain across surviving pool
- Return AgentResponse(recommended_asins[:10], ask_attribute)
8. Benchmark Results
Shopping Copilot was evaluated across all 200 public evaluation sessions against the starter baseline:
| Metric | Starter Baseline | Shopping Copilot | Delta |
|---|---|---|---|
| Hit Rate@10 | 12.5% | 98.0% | +85.5% |
| Mean Reciprocal Rank (MRR) | 0.068 | 0.667 | +0.599 |
| Mean Turns to Target (MTTC) | 9.81 turns | 2.41 turns | -7.40 turns |
| Efficiency Metric | 0.125 | 0.859 | +0.734 |
| Recommended Technical Score | 0.107 | 0.862 | +0.755 |
| Turn Latency | ~20 ms | ~42 ms | Real-time interactive |
| API Cost | \$0.00 | \$0.00 | Zero cost |
Per-Scenario Breakdown
Scenario: BUYING (80 sessions)
- Hit Rate@10: 98.75% (Baseline: 22.5%)
- MRR: 0.6461
- MTTC: 1.88 turns
Scenario: BROWSING (80 sessions)
- Hit Rate@10: 97.50% (Baseline: 2.5%)
- MRR: 0.6418
- MTTC: 2.33 turns
Scenario: INTENT OVERRIDE (30 sessions)
- Hit Rate@10: 96.67% (Baseline: 13.3%)
- MRR: 0.7520
- MTTC: 4.03 turns
Scenario: BOUNDARY / FALLBACK (10 sessions)
- Hit Rate@10: 100.0% (Baseline: 0.0%)
- MRR: 0.7833
- MTTC: 2.50 turns
9. Experimental Ablation Study
Progress was guided by a 24-run ablation log, modifying one variable at a time and evaluating the full 200-session benchmark before committing:
| Run | Hypothesis Tested | HR@10 | MRR | MTTC | Decision and Rationale |
|---|---|---|---|---|---|
| 0 | Starter Baseline | 12.5% | 0.068 | 9.81 | Initial baseline measurement |
| 1 | Title Deduplication (>85% similarity) | 82.0% | 0.433 | 4.40 | KEEP: Freed redundant recommendation slots |
| 2 | Hard Facet Filtering (Color/Material/Size/Budget) | 82.5% | 0.444 | 4.40 | KEEP: Eliminated category and attribute mismatches |
| 5 | Recency Weighting (1.0 + N * 0.2) | 82.5% | 0.444 | 4.42 | KEEP: Prioritized recent user messages |
| 12 | Category Slot Extraction from Opener | 84.5% | 0.480 | 4.23 | KEEP: Anchored broad browsing queries |
| 14 | Intent Override: Reset shown_asins on Mind-Change |
96.5% | 0.552 | 3.43 | KEEP: Fixed pre-flip target lockout bug |
| 16 | Bigram Phrase Match Bonus (+5.0 verbatim) | 97.0% | 0.613 | 2.89 | KEEP: Rewarded multi-word product specifications |
| 19 | Category Match-Count Gradient Bonus | 97.0% | 0.635 | 2.85 | KEEP: Replaced flat bonus with matching gradient |
| 22 | Vocabulary Mining (Jewelry metals, fabrics, colors) | 98.0% | 0.642 | 2.83 | KEEP: Filled metadata vocabulary gaps |
| 23 | Information Gain Question Policy in _pick_attribute |
98.0% | 0.657 | 2.38 | KEEP: Reduced MTTC from 2.83 to 2.38 turns |
| 24 | Title Recency Reranking + Brand Store Anchor | 98.0% | 0.667 | 2.41 | KEEP: Reached peak MRR and 0.862 score |
10. Practical Impact for E-Commerce Platforms
- Infrastructure Cost Savings: At 10 million daily active users averaging 3 turns each, an LLM-based search agent generates 30 million API calls per day (\$150,000 to \$300,000 per month). Shopping Copilot runs in-memory on standard server instances with zero API bills.
- Search Drop-Off Reduction: Reducing average turns to discovery from 9.81 to 2.41 turns decreases search fatigue and shopping cart abandonment.
- Live Stream Integration: With turn latencies under 45 ms, Shopping Copilot can run as an interactive search layer in live streaming and mobile video apps without lag.
11. Engineering Challenges and Solutions
Intent Override Pre-Flip Lockout (Run 14):
- Issue: In intent override sessions, users change preferences at turn 3 or 4. Our Guess-Who mechanism recorded every product shown in turns 1 and 2 in
shown_asins. When the user switched intents, the newly targeted item was often blacklisted if it had appeared earlier as an irrelevant candidate, causing Intent Override Hit Rate to stall at 13.3%. - Fix: Added an override listener that detects reversal triggers, clears
shown_asinsfor the new category context, and flushes conflicting slots. This lifted Intent Override Hit Rate from 13.3% to 96.7%.
- Issue: In intent override sessions, users change preferences at turn 3 or 4. Our Guess-Who mechanism recorded every product shown in turns 1 and 2 in
The Clean Slate Problem (Run 8 vs. Run 12):
- Issue: When addressing intent overrides in Run 8, we initially wiped all conversation terms on override. Accuracy dropped from 16.7% to 13.3%.
- Root Cause: The opening turn ("I'm looking for sandals...") contained the category anchor that narrowed 50,000 items to ~300. Wiping all terms removed that anchor, leaving generic modifiers ("leather") to match irrelevant apparel.
- Fix: Separated persistent structural slots (
category,user_profile) from volatile modifier slots (color,material,size,budget), resetting only the conflicting modifier slots.
Near-Duplicate Listing Flooding (Run 1):
- Issue: Catalogs often contain multiple listings of the same product with minor title variations. A top-10 list would return 10 identical t-shirts in different sizes, wasting 9 recommendation slots.
- Fix: Added
difflib.SequenceMatcherto measure title string distance and filter items with >85% similarity against higher-ranked candidates. This raised Hit Rate from 12.5% to 82.0%.
Phrase Syntax vs. Bag-of-Words Noise (Run 6 vs. Run 16 & 17):
- Issue: Enforcing strict boolean phrase matching in FTS5 queries caused Hit Rate to drop from 82.5% to 77.5% because natural conversation rarely matches catalog wording verbatim.
- Fix: Used broad OR-unions for candidate retrieval (gathering the top 300 candidates) and applied an in-memory n-gram reranking bonus (+5.0 bigrams, +7.0 trigrams) during ranking.
Sparse Attribute Metadata (Run 22):
- Issue: Jewelry and accessories often embed key attributes ("sterling silver", "rose gold", "chiffon") in descriptions rather than structured dictionaries.
- Fix: Mined catalog text across all 50,000 items to build broader regex vocabularies for materials, metals, and colors, increasing Browsing Hit Rate to 97.5%.
Clarification Policy and Wasted Turns (Run 23):
- Issue: Naive question selection asks about attributes with no variation in the active candidate pool or continues asking questions when the pool has already shrunk to fewer than 5 items.
- Fix: Implemented Shannon entropy scoring over surviving candidate attributes and added an endgame rule: when pool size is $\le 10$, the agent stops asking questions and directly returns the candidate list.
Determinism and Crash Safety (Run 21):
- Issue: Random tie-breaking and iteration over unordered sets caused minor score variations (+/-0.0003) between test runs.
- Fix: Replaced randomized sorting with deterministic tie-breaking on
parent_asinand converted sets to sorted tuples, ensuring repeatable results and crash-safe execution.
12. Accomplishments
- Improved Hit Rate@10 from 12.5% to 98.0% on a 50,000-product catalog.
- Reduced mean turns to target from 9.81 to 2.41 turns.
- Achieved sub-45 ms response times with zero external API dependencies ($0.00 cost).
- Maintained 100% test pass rate across unit and adversarial test suites.
13. Future Roadmap
- Local Dense Vector Retrieval: Integrate a quantized on-device embedding model (such as
bge-smallvia ONNX) to supplement BM25 on abstract aesthetic searches ("cottagecore aesthetic outfit"). - Expanded Attribute Elicitation: Expand entropy-based question selection to cover brand, silhouette, sleeve length, and price tiers.
- Visual Search Input: Allow users to submit image references or video frames to seed the session state.
14. Judging Criteria Alignment
| Judging Criterion | Weight | Project Implementation |
|---|---|---|
| Technical Execution | 35% | In-memory SQLite FTS5 search index, multi-stage filtering, entropy-based question selection, and 100% test pass rate on adversarial suites. |
| Innovation and Insight | 20% | Formulated conversational search as active information-theory elimination with Guess-Who deduplication and intent override recovery. |
| Impact and Relevance | 20% | Directly addresses conversational commerce by reducing turns to purchase from 9.8 to 2.4 and cutting search bounce rates. |
| Feasibility and Practicality | 15% | Standard-library implementation with zero API costs, 42 ms turn latency, and no external service dependencies. |
| Presentation and Documentation | 10% | Complete documentation including REPORT.md, README.md, DEVPOST_SUBMISSION.md, and reproducible experiment logs. |
15. Reproducibility
To run the complete benchmark evaluation locally:
# Clone the repository
git clone https://github.com/Bread04/TIKTOKJAM-.git
cd TIKTOKJAM-
# Run the evaluation benchmark (Python 3.10+ standard library, no pip install needed)
python -m evaluator.local_evaluator
Challenges we ran into
The Pre-Flip Intent Override Lockout Trap:
- The Problem: In our early iterations, Intent Override hit rate was severely lagging at 20.0% (compared to >90% across buying and browsing). Our initial theory was that pre-flip preference keywords were lingering in the query buffer and polluting post-flip search. However, wiping those terms actually dropped the override hit rate further to 13.3%.
- The Discovery: Replaying session traces via our diagnostic script (
scripts/debug_override.py) revealed that in 22 of 30 override sessions, our agent was already ranking the true target into the Top 10 before the evaluator script fired the mind-change at Turn 3 or 4. Because hits were disabled before the flip, our "Guess Who" candidate-elimination state (shown_asins) recorded the target as a rejected product and permanently barred it from being shown again. - The Fix: We built a stateful intent-override detector that selectively flushes
shown_asinswhile preserving category slot anchors upon detecting a change of mind, immediately lifting Override Hit Rate from 20.0% to 96.7%.
Balancing Hard Facet Pruning vs. Semantic Recall in 50,000 Products:
- In a frozen 50,000-product catalog with sparse, unstructured text, treating attribute constraints as soft ranking bonuses often allowed irrelevant items with high term frequencies (e.g., women's winter boots on a men's shoe query) to outrank actual matches.
- Conversely, rigid SQL filtering risked emptying the candidate pool on ambiguous or multi-token phrasing.
- We resolved this by building a hybrid filter layer: strict facet pruning for high-confidence attributes (color, material, size, and budget with a $\pm 30\%$ tolerance band) paired with a safety-valved category filter (enforced only when matches constitute $\ge 5$ items or $\ge 5\%$ of the active candidate pool).
Candidate Pool Width vs. Ranking Noise:
- Widening the SQLite FTS5 candidate retrieval pool from 300 to 500 items in an attempt to catch edge-case misses introduced significant ranking noise and degraded Hit Rate@10 by 3.5%.
- We learned that retrieval width must remain tightly bounded ($k=300$), and that exact multi-word phrase bonuses (bigrams and trigrams) and category match-count gradients yielded far cleaner candidate separation than simply fetching a wider pool.
Multi-Turn Determinism & Replay Fidelity:
- Non-deterministic behaviors (such as SQLite
ORDER BY RANDOM()fallback padding and unordered Pythonsetiteration) caused scoring fluctuations of $\pm 0.03\%$. - We enforced 100% byte-identical reproducibility across benchmark runs by implementing deterministic
ORDER BY parent_asintie-breaking, ordered attribute tuples, and comprehensive try/except error boundaries on every turn.
- Non-deterministic behaviors (such as SQLite
Accomplishments that we're proud of
Near-Perfect Benchmark Performance Across All 200 Test Sessions:
- Hit Rate@10: Jumped from 12.5% (baseline) to 98.0% (+85.5% absolute gain).
- Mean Reciprocal Rank (MRR): Rose from 0.068 to 0.667 (a nearly $10\times$ improvement in ranking precision).
- Mean Turns to Target (MTTC): Dropped from 9.81 turns down to 2.41 turns (median of ~2.0 turns to reach the target item).
- Overall Technical Score: Increased from 0.107 to 0.862 (+0.755).
Dominating Browsing & Override Scenarios:
- In vague Browsing sessions (where users start with ambiguous ideas), our hit rate surged from 2.5% to 97.5%.
- In complex Intent Override sessions, our recovery hit rate reached 96.7% with an MRR of 0.752.
- On Boundary sessions, our agent achieved a flawless 100% Hit Rate@10 and 0.783 MRR.
$0.00 Token Cost & Sub-45ms Per-Turn Latency:
- Engineered entirely in pure Python using only standard libraries and an in-memory SQLite FTS5 BM25 search index.
- Zero external paid LLM API calls, zero token costs, 100% data privacy, and the ability to evaluate all 200 multi-turn sessions in under 84 seconds on standard CPU hardware (~420 ms per full 10-turn dialogue).
Disciplined, Evidence-Driven Engineering:
- Documented 24 distinct, logged ablation experiments in our lab notebook (
docs/experiments/experiments.md), adhering strictly to a single-variable change protocol to isolate every architectural gain.
- Documented 24 distinct, logged ablation experiments in our lab notebook (
What we learned
- Conversational Search is an Information Game ("Guess Who?"), Not an Isolated Search Box:
- The stock baseline failed because it treated every turn as an independent query. By maintaining state across turns—tracking extracted slots, recency-weighted term history, user profile preference tags, and previously displayed ASINs—we could systematically eliminate non-target items from the candidate space.
- Clarifying Questions Must Maximize Information Gain (Entropy Reduction):
- Asking generic questions wastes valuable turns. By implementing an adaptive question policy (
ask_attribute: "other"and candidate-diversity evaluation across color, material, and size), the agent dynamically bisects the remaining candidate pool and extracts ground-truth product attributes directly from the user's natural language responses.
- Asking generic questions wastes valuable turns. By implementing an adaptive question policy (
- Signals Must Have Gradients, Not Just Binary Offsets:
- Moving from flat boolean bonuses (e.g., $+5.0$ if any category word matched) to a match-count gradient ($+3.0 \times \text{words matched}$) plus exact bigram/trigram phrase boosts yielded significant MRR jumps (+0.061 to +0.080) by pulling genuine syntactic matches to the very top rank.
- Structural Completeness Beats Parameter Tuning:
- Hyperparameter sweeps over BM25 column weights yielded negligible impact until we noticed missing slot signals (like missing category extraction from
"looking for X"openers). Inspecting dead logic paths and rule integrity consistently outperformed brute-force grid search.
- Hyperparameter sweeps over BM25 column weights yielded negligible impact until we noticed missing slot signals (like missing category extraction from
What's next for Shopping Copilot
- Hybrid Dense + Sparse Semantic Retrieval:
- Integrate lightweight, local ONNX/quantized embeddings (such as MiniLM or BGE-small) alongside our BM25 FTS5 index to support semantic matching for conceptual browsing queries (e.g., "something breathable for a summer beach wedding" or "cozy aesthetic streetwear") without requiring exact keyword matches.
- Full-Spectrum Entropy Question Selection:
- Expand the active attribute diversity calculation beyond color, size, and material to include brand store anchors, style archetypes, price tiers, and specific functional features, maximizing information gain on highly dispersed candidate pools.
- Adaptive Real-Time Budget & Sentiment Parsing:
- Replace regex-based budget parsing with comparative semantic range estimators capable of handling colloquial price constraints (e.g., "under budget", "luxury grade", "affordable alternative", "best bang for buck").
- Live TikTok Shop Interactive UI & Live-Stream Integration:
- Package Shopping Copilot into a drop-in WebSocket microservice for live stream commerce, empowering TikTok creators and automated shopping assistants to recommend real-time pinned products tailored to audience chat questions with zero latency.

Log in or sign up for Devpost to join the conversation.