Inspiration
Traditional e-commerce keyword search frequently fails when consumers express vague, evolving, or contradictory shopping intent. Standard platforms treat conversations as static query appends, making it difficult to narrow down a 50,000-item catalog without exhausting the user. We set out to build a lightweight, ultra-fast shopping copilot that proactively asks clarifying questions, updates user intent dynamically, and places target items in the Top 10 with zero paid API dependencies.
What it does
Our agent (FastAgent) acts as an intelligent backend shopping assistant for the Amazon Clothing, Shoes, and Jewelry catalog. It:
- Parses Customer Intent: Extracts key attributes and hard constraints while filtering out filler words.
- Maintains Dynamic State: Updates active slots incrementally and gracefully handles intent overrides (e.g., clearing old slots when a user changes their mind mid-conversation).
- Proactively Clarifies: Chooses optimal missing attributes to ask the user, rapidly shrinking the search space.
- Ranks & Retrieves: Combines category locking, exact constraints, lexical token matching, and popularity fallbacks to deliver a precise Top 10 product recommendation payload (
parent_asin) in under 3 turns.
How we built it
We engineered a modular Python pipeline designed to run 100% offline and CPU-only:
agent/fast_agent.py&agent/state.py: The primary brain and session memory tracker managing slot accumulation and overrides.agent/parsing.py: Independent pattern extractor designed to handle filler words, clause reordering, and complex phrasing.agent/question.py: Strategic decision engine for triggering attribute clarification prompts.agent/catalog.py&agent/routes/: High-speed in-memory indexing for category filtering, constraint matching, and lexical retrieval.ui/server.py&ui/static/: A lightweight browser-based UI for visual testing and demonstration.
Challenges we ran into
Our biggest hurdle was maintaining parsing accuracy when users reworded input utterances, reordered clauses, or added filler text. Order-dependent regex matching initially dropped performance by 45.16% on reworded queries. We resolved this by decoupling event extraction, stripping filler tokens, and executing independent pattern searches across the full utterance while preserving the anchored path for clean inputs.
Accomplishments that we're proud of
- High Technical Score: Achieved a local evaluation score of 0.852704 ($\text{HitRate@10} = 0.9600$, $\text{MRR} = 0.6813$, $\text{MTTC} = 2.585$).
- Zero Cost & High Efficiency: Runs completely offline with 0 tokens used and 0 paid API calls.
- Proven Robustness: Demonstrated a minimal performance drop (1.01% gap) under aggressive paraphrasing and reworded test sets.
What we learned
We learned that deterministic, well-engineered rule systems and dynamic state tracking can outperform non-deterministic LLM pipelines in constrained search environments—offering vastly superior speed, zero runtime cost, and total reliability.
What's next for kpopy demon hunter - Shopping Copilot
- Integrating lightweight local semantic embeddings (e.g., Model2Vec) to improve contextual similarity on long-tail product descriptions.
- Expanding the dynamic state machine to handle multi-category intent switching within a single session.
Built With
- bm25s
- ci-cd
- conversational-ai
- css3
- git
- html5
- information-retrieval
- javascript
- pytest
- python
- recommender-systems
- regex
- vscode
Log in or sign up for Devpost to join the conversation.