Threadline - Say it Once
Threadline finds the right product in the top 10 results 87.5% of the time in about 3.4 turns of conversation, with zero LLM calls. Here's how we got there, including the ideas that didn't work.
Inspiration
Most conversational shopping search treats every message like the first one. Say "I need running shoes," get asked a question, answer it, and the system has already half-forgotten your original request. Say "actually, never mind the color, I need it waterproof instead," and most systems either ignore the correction or get confused trying to reconcile it.
We wanted to build something that treats a shopping conversation like an actual conversation: constraints accumulate, intents shift, sometimes a shopper contradicts themselves, and the system should get better at helping the longer you talk to it, not more confused.
What does it do
Threadline is a conversational shopping agent, built for the TechJam 2026 Shopping Copilot challenge, that:
- Remembers the whole conversation. Every turn's revealed constraints whether its material, color, budget, use case accumulate into a running picture of what the shopper wants, instead of each turn starting from zero.
- Detects buying vs. browsing intent from how the shopper phrases things, and adapts its retrieval accordingly, locking onto a stated hard requirement when someone's clearly buying, staying broad and diverse when someone's just exploring.
- Handles contradictions gracefully. When a shopper says "actually, ignore that, I need X instead," Threadline doesn't get stuck averaging old and new preferences into a confused middle ground. It adapts and keeps moving.
- Asks smart, efficient questions. Instead of working through a fixed checklist of attributes (some of which a given product may not even have, wasting a turn), it leads with an open elicitation strategy. It only falls back to specific questions when needed, cutting wasted turns significantly.
- Runs entirely locally - no LLM calls, no external API, no cost per request. Retrieval and ranking are built on SQLite's FTS5 full-text search with BM25, plus lightweight rule-based logic.
Result on the organizer's 200-session public evaluation set: 87.5% Hit Rate@10, an average of ~3.4 turns to find the right product, and a TechnicalScore of 0.758.
How we built it
The foundation is a stateful session tracker that sits atop an in-memory SQLite FTS5 index for the 50,000-product catalog. Each turn's newly revealed constraints get folded into the running session state; retrieval runs a BM25-ranked search over everything accumulated so far, then re-scores candidates by how many distinct turns' worth of constraints they actually satisfy - rewarding well-matched products without ever hard-excluding one over a single missing term.
On top of that:
- A lightweight regex-based intent classifier routes sessions onto a buying or browsing track, re-evaluating every turn so a session can switch tracks mid-conversation.
- A fallback question ladder re-orders itself at runtime based on which attributes the current candidate pool is actually likely to have answers about, rather than always asking in the same fixed order.
- Filler-response detection and override detection (regex-driven) let the agent tell a real answer apart from a scripted non-answer, and recognize when a shopper is contradicting an earlier statement.
Mapped to the challenge's four pillars, honestly:
- Core Architecture - Intent Routing & Hybrid Pipeline: covered via the buying/browsing classifier and BM25 + coverage-boost hybrid ranking. We deliberately did not add an LLM semantic-ranking stage (see below).
- Dialog Strategy - Multi-Turn Scenario Evolution: covered via the session state tracker, incremental slot accumulation, and override handling.
- Self-Evolution - Dynamic Context Programming: partially covered by the runtime-adaptive question ordering, which is real and measured; long-term cross-session user profiles are implemented and functional but untestable on this dataset (every evaluation session is a distinct user, per TechJam organizer's own docs).
- Evaluation Matrix: addressed directly that every design decision below was validated against the organizer's local evaluator before being kept or discarded.
Every change was tested against the organizer's local evaluator (Hit Rate@10, MRR, MTTC) before being kept. Two features we built and tested were kept out of the final agent because the data said no:
- Seeding sessions with a
preference_tagsfield from the user profile. Sounds like free personalization — measured, it dropped TechnicalScore from 0.758 to 0.735 (hit rate 0.875 → 0.850). The tag values ("fit," "comfort," "durability") turned out to be generic marketing language present across most of the catalog, so seeding them just added noise to retrieval. - A local TF-IDF vector-similarity re-ranking stage, as a lightweight, no-cost stand-in for the "LLM Semantic Ranking" the pillar describes. Measured, it dropped TechnicalScore from 0.757 to 0.565 (hit rate 0.875 → 0.665) — worse across every scenario, not an edge case. BM25 already does field-weighting and term-frequency saturation on the same candidate pool; a flat, unweighted cosine similarity turned out to be a strictly blunter signal, and letting it override BM25's already-good ordering hurt rather than helped.
Both remain in the codebase, disabled behind feature flags with the measured numbers documented in-line, a deliberate choice not to delete tested negative results.
A note on why there's no LLM here. Solutions with access to OpenAI or Google Cloud Vertex APIs can bring genuine semantic understanding to retrieval that a BM25-based approach doesn't have. We chose a different trade-off deliberately: zero external cost, zero per-request latency, and a fully deterministic result — the same conversation produces the same output every time, with no API dependency that could fail or rate-limit during judging. To be precise about what we did and didn't test: the "vector similarity" stage we built and rejected was a lightweight local TF-IDF cosine similarity, not a genuine embedding model — a real embedding-based semantic layer might well outperform BM25 here. We didn't test that, because doing so would reintroduce the external dependency we chose to avoid. So this isn't a claim that AI doesn't help — it's a claim that a zero-cost, fully local, reproducible system reaching 87.5% top-10 accuracy was the trade-off we wanted to make for this challenge, where the rules explicitly state a paid LLM isn't required to compete.
Challenges we ran into
- A silent dead-code bug. Early on, a duplicate method definition meant our more sophisticated retrieval logic was never actually running — the agent was silently falling back to a much simpler search the whole time. Found by tracing why a redesign had zero effect on scores.
- Wasted turns were quietly costing us efficiency. Our first working version asked about attributes in a fixed order — material, then color, then size, and so on — even when a product had nothing to say about a given attribute, burning a whole turn for zero new information. Switching to an open-ended elicitation strategy first, and only falling back to the fixed list once that's genuinely exhausted, cut mean turns-to-conversion from 4.33 to 3.59 and lifted the efficiency component of our score from 0.667 to 0.741 — TechnicalScore moving from 0.726 to 0.740 from that one change alone.
- Hard filtering looked right and measurably wasn't. We tried requiring every conversation turn's words to literally co-occur in a product's listing (a strict AND across turns). It hurt hit rate — structured info like a stated budget rarely appears as literal prose in a title or description, so a hard filter can exclude the correct item outright. We replaced it with a soft, coverage-based boost that rewards multi-turn matches without ever excluding a candidate.
- A plausible fix that the data proved wrong. When a shopper contradicts an earlier preference, our instinct was to reset the conversation's accumulated context back to a clean baseline before adding the new value — otherwise, we reasoned, the old contradicted value would pollute the query. Tested against real evaluation data, this was worse: resetting also throws away other still-valid info gathered along the way (budget, brand, etc.), and it turned out the old value doesn't meaningfully hurt (matching logic is inclusive, not exclusionary). Not resetting scored significantly better on override-heavy conversations. We reverted the "fix" once the numbers said so.
- A real performance bug at catalog scale. A scoring step that issued one database query per accumulated conversation turn worked fine on a small test catalog but became a serious bottleneck at 50,000 products across hundreds of evaluation sessions. Rewritten to fetch each candidate's data once per turn and score in memory; a large, measured speedup with no change in behavior.
Accomplishments that we're proud of
- A fully working, evaluation-validated conversational agent with zero external API dependency, deterministic, reproducible, and free to run.
- A genuine before/after data trail for every major design decision, including the ones that didn't work; we think that's a more honest and more useful account of the engineering than a story with no wrong turns in it.
- Handling all four scenario types the challenge defines (buying, browsing, no-stated-preference, and intent-override) with a single coherent architecture.
What we learned
The biggest lesson: test the "obviously correct" idea before trusting it. More than once, a design that reasoned soundly step-by-step underperformed a simpler alternative once measured on real data, and the fix was to trust the evaluator over our own intuition. We also learned to be honest about scope: we deliberately did not build an LLM-based semantic ranking stage or a hard binary intent classifier the original problem statement describes, because a local, rule-based alternative was faster, free, and when we tested a TF-IDF vector-similarity stage as a substitute, measurably worse on this dataset. Knowing when not to add a fashionable-sounding component was as valuable as anything we built.
What's next
- A genuine long-term user-profile mechanism - the interface supports it, but the competition's dataset never repeats a user, so we couldn't validate it here. Rather, a real multi-session deployment would.
- A learned (rather than heuristic) intent classifier for the buying/browsing split, trained on labeled conversation data if it were available.
- Further investigation into why
intent_overrideconversations still take longer on average than other scenarios, likely a retrieval-recall question worth targeted testing.
Log in or sign up for Devpost to join the conversation.