Inspiration

Shopping searches rarely arrive as complete specifications. A customer may start with “I need shoes for a trip,” reveal a budget later, or replace an earlier preference entirely. A useful shopping agent must decide what to ask, remember what still matters, forget what has been superseded, and rank the right product early—not merely append every word to one search query.

What it does

Shopping Copilot is one offline Python Agent for the organizer's frozen 50,000-product Clothing, Shoes and Jewelry catalog. During each simulated conversation it maintains isolated session state, asks a structured clarifying question, and returns up to ten unique catalog-valid parent_asin identifiers in ranked order. It treats vague browsing, newly revealed constraints, repeated turns, and changed intent differently. When the shopper changes direction, obsolete requirements are removed from active retrieval and the shown-product page is reset.

How we built it

At startup the Agent constructs a process-private SQLite FTS5 index from the organizer-provided catalog. It retrieves through strict and progressively relaxed Unicode lexical routes, then reranks with visible catalog evidence, active-constraint coverage, price compatibility, and a weak fixed popularity prior. After a miss it rotates to a different valid page instead of repeating the same list.

The runtime is Python 3.9+ using only the standard library and SQLite FTS5. It requires no network access, API key, hosted service, model weight, or third-party Python package. Temporary-index failures have a bounded in-memory retry. The same catalog and conversation produce the same answer, while each new turn can still change state, retrieval, ranking, and pagination.

Testing and evidence

The organizer's BM25 reference reports Hit Rate@10 0.125000, MRR 0.068034, MTTC 9.810000, and TechnicalScore 0.106710 on the released 200-session development evaluator. One frozen exact-package observation for this Agent reported Hit Rate@10 0.990000, MRR 0.630409, MTTC 2.030000, and TechnicalScore 0.863523. These public sessions influenced development, so that result is reported as contaminated development evidence—not as a prediction of the 800 private sessions.

The strongest less-contaminated Track-specific evidence is a precommitted, one-time reserve of 140 train-only targets disjoint from upstream validation, test, and an earlier audit. The exact pre-storage ranking lineage reached Hit Rate@10 0.978571, MRR 0.678245, MTTC 2.457143, and TechnicalScore 0.863616. Intent Override was weaker at 18/21. The current runtime adds output-neutral parser/storage hardening and was deliberately not rerun on that consumed reserve. This evidence still shares the released catalog and simulator and is not a private-score estimate.

The repository includes contract, state, packaging, failure-injection, contamination, provenance, and fuzz tests. The final local suite discovered 88 tests: 80 executed and passed, eight research-only NumPy tests were intentionally skipped, and none failed. Additional frozen gates cover 800,000 target-free protocol assertions and 2,400 interleaved session calls with zero exceptions or contract violations.

What we learned

Adding a general-purpose model was not automatically better. Frozen semantic and API reranking experiments sometimes improved their native benchmarks or an overall aggregate, but harmed Track-specific transfer, Intent Override, or Boundary behavior. We retained the smaller offline Agent because no tested model passed every predeclared promotion gate. This is a task-specific result, not a claim that neural retrieval is generally inferior.

Challenges and limitations

The private score remains unknown. Lexical retrieval is less robust to unrestricted real-world paraphrases, aggregate profile fields are ignored after a failed validation gate, some ties depend on the organizer's frozen catalog row order, and SQLite FTS5 must be present in the evaluation Python build. Local timing and memory measurements do not guarantee another host. No revenue, conversion, satisfaction, production deployment, or final-session impact is claimed.

Contribution and disclosure

Shreyansh Agarwal is the entrant and, if submitted solo, the authorized representative. He defined the project objective, supplied and reviewed the official challenge material, directed the evaluation and risk boundaries, reviewed design decisions, authorized experiments, and selected the frozen submission candidate. OpenAI Codex assisted with source inspection, implementation, experiments, testing, packaging, and documentation under the entrant's direction. The submitted runtime itself uses no OpenAI API or model.

Built With

  • conversational-search
  • fts5
  • information-retrieval
  • python
  • sqlite
  • testing
  • unicode
  • unittest
Share this project:

Updates

Submission history