Inspiration

Shopping is rarely a single, static query. Requirements accumulate, disappear, and sometimes reverse completely:

“I want black running shoes” → “Actually, not black” → “Show me casual shoes instead.”

Many shopping assistants bury these changes inside an ever-growing prompt. ConstraintGraph takes a different approach: it represents conversational intent as explicit, inspectable state and treats product search as progressive uncertainty reduction.

What it does

ConstraintGraph is an adaptive shopping agent built for the TechJam 2026 Shopping Copilot challenge. It searches the official 50,000-product catalog and returns up to ten valid recommendations on every turn.

Its core capabilities include:

  • Event-sourced conversational state using explicit ADD, REMOVE, NO_PREFERENCE, and RESET events
  • Separate pipelines for precise Buying requests and exploratory Browsing
  • Exact constraint intersection for high-precision retrieval
  • Broad, profile-aware candidate generation for discovery
  • Information-gain clarification questions
  • Adaptive BM25 and word/character TF-IDF after ambiguous intent resets
  • Intent-override handling that prevents obsolete preferences from leaking into later recommendations

For each unanswered attribute (A), the agent estimates how much asking about it would reduce the candidate pool:

$$ IG(A) = H(C) - \sum_{v \in A} P(v)H(C \mid v) $$

It then asks the question expected to reduce uncertainty the most.

The entire system runs locally on CPU with zero runtime LLM calls, API calls, prompt tokens, completion tokens, or model cost.

How I built it

ConstraintGraph is a deterministic Python retrieval and orchestration system.

The conversation layer records every preference change in an append-only event log. A reducer replays those events to produce the current intent, making state changes easy to inspect and test.

The routing layer classifies the current state as either:

  • Buying: explicit constraints, precision-first retrieval, and exact product-signature matching
  • Browsing: broad retrieval, controlled diversity, and clarification before transitioning toward a purchase

The retrieval layer combines:

  • normalized reverse postings over category, color, material, price, features, and product details;
  • SQLite FTS5 BM25;
  • sparse word and character TF-IDF with scikit-learn;
  • route-specific ranking and controlled diversity.

Exact matching remains dominant during ordinary Buying interactions. Lexical fusion activates selectively after an intent reset, where a replacement clue may be ambiguous.

The implementation uses only participant-visible catalog fields. No target product identifiers are hard-coded, the catalog remains read-only, and the official evaluator is unchanged.

Challenges I faced

Handling intent overrides

The hardest state-management problem was deciding what should survive a user correction. A complete reset can incorrectly discard a still-valid category, while a partial update can leave stale colors, brands, or features behind.

I solved this with explicit intent generations and preference-scoped resets. Constraints and previous clarification questions are cleared while a compatible category can remain active.

Balancing exact and lexical retrieval

Applying BM25 and TF-IDF to every Buying request initially seemed attractive, but ablation testing showed that unrestricted lexical fusion could weaken exact-match precision.

The final architecture therefore activates lexical fusion only when ambiguity warrants it—particularly after an intent reset—while retaining exact constraint retrieval for ordinary Buying requests.

Choosing complexity based on evidence

I initially expected semantic embeddings or an LLM retriever to be necessary. However, the frozen deterministic architecture already achieved 0.975 Hit Rate@10 on the pseudo-hidden split.

Adding semantic infrastructure would have increased dependency, latency, hardware, and regression risk without measured evidence of better official metrics. I therefore kept the final system deterministic, reproducible, and locally executable.

Results

On the official 200-session public evaluator, ConstraintGraph achieved:

Metric Official starter ConstraintGraph
Hit Rate@10 0.125 0.995
MRR 0.0680 0.7308
MTTC—lower is better 9.81 2.19
Efficiency 0.119 0.881
TechnicalScore 0.1067 0.8929
Runtime LLM tokens 0 0

It successfully retrieved the target within the top ten for 199 of 200 public sessions and ranked 127 targets first.

A fixed 40-session pseudo-hidden split achieved 0.975 Hit Rate@10 and a 0.8703 TechnicalScore. These are public-development results and do not guarantee performance on the final evaluation set.

What I learned

The main lesson was that an effective shopping agent needs intelligent orchestration more than complexity for its own sake.

Explicit state makes corrections reliable. Route-aware retrieval prevents exploratory browsing and exact purchasing from weakening each other. Information theory provides a principled way to choose clarification questions. Most importantly, controlled ablations make it possible to reject additional complexity when the evidence does not justify it.

I also learned that reproducibility is a product feature. ConstraintGraph can explain its state transitions, rebuild its indexes from visible catalog data, run without external services, and report exactly where each recommendation came from.

What’s next

Future work includes:

  • testing against genuinely free-form human shopping language;
  • calibrating clarification answerability on a larger conversation corpus;
  • evaluating a lightweight semantic retriever on a separately frozen protocol;
  • learning fusion weights without using competition-specific target exceptions;
  • reducing cached startup time through more portable precomputed indexes.

ConstraintGraph shows that transparent state, adaptive retrieval, and carefully measured engineering can turn a changing conversation into an efficient path toward the right product.

Built With

Share this project:

Updates

Submission history