TikiTaka — Conversational Shopping Copilot

How our solution addresses the problem

TikiTaka is a conversational shopping agent designed to identify a customer’s intended product from a frozen catalog of 50,000 items within 10 turns.

Instead of treating each message as an isolated search query, TikiTaka maintains structured session state containing the shopper’s preferences, exclusions, corrections, and intent changes. It distinguishes between browsing and purchase-focused behaviour and chooses one action per turn:

  • Ask a high-value clarification question when the answer is likely to improve the ranking.
  • Recommend up to 10 catalog-valid products when enough information is available.

Candidate products are retrieved using BM25 text search, optional dense embeddings, and structured product metadata. The results are fused and reranked using the shopper’s active constraints. Deterministic validation ensures that model output cannot introduce invalid product IDs, unsupported attributes, or unsafe state changes.

The submitted configuration also includes a deterministic, network-free fallback, allowing the agent to remain functional when API credentials or internet access are unavailable.

Development tools used

  • Python 3.10+ — primary development and runtime language.
  • Git and GitHub — source control, collaboration, and milestone-based reviews.
  • OpenAI Codex and Claude Code — AI-assisted development and code review.
  • Python unittest — automated unit, integration, fault-injection, and contract testing.
  • SQLite FTS5 — in-memory full-text indexing and BM25 retrieval.
  • Official local evaluator — deterministic evaluation across the 200 public sessions.
  • Custom build and verification scripts — submission packaging, checksum validation, dependency auditing, and network-disabled testing.

APIs used

  • OpenAI Chat Completions API

    • Model: gpt-5.6-terra
    • Reasoning effort: medium
    • Used for intent interpretation, structured state extraction, query rewriting, clarification phrasing, and shortlist reranking.
    • Optional at runtime because the agent can fall back to deterministic logic.
  • OpenAI Embeddings API

    • Model: text-embedding-3-large
    • Embedding size: 1,024 dimensions
    • Used to build and evaluate a dense semantic index for the product catalog.
    • The final submitted route does not require the generated dense index.

No other external APIs are required.

Libraries and frameworks used

The submitted agent has no mandatory third-party runtime dependencies. It runs primarily on the Python standard library, including:

  • sqlite3 — FTS5 indexing and BM25 retrieval.
  • urllib — direct HTTPS communication with the OpenAI APIs.
  • array and math — vector storage and cosine similarity.
  • dataclasses and typing — structured domain contracts.
  • json and hashlib — data loading, manifests, and checksum validation.
  • pathlib, collections, decimal, re, and unicodedata — file handling, normalization, scoring, and cost accounting.

Optional development-only libraries include:

  • NumPy — accelerated exact dense-vector search and index construction.
  • tiktoken — token counting and API cost estimation.

We did not use PyTorch, TensorFlow, Hugging Face Transformers, or a model-training framework. No foundation model was fine-tuned.

Datasets and assets used

  • Amazon Reviews 2023 — McAuley Lab, UCSD

    • Competition data is derived from the Clothing_Shoes_and_Jewelry category.
    • Products are identified using their Amazon parent_asin.
    • Product data includes titles, categories, features, descriptions, prices, ratings, stores, and structured details.
  • Frozen product catalog

    • Contains 50,000 products.
    • Treated as read-only.
    • Only textual and structured metadata is used.
  • Public competition sessions

    • Contains 200 labeled development sessions.
    • Covers Buying, Browsing, Intent Override, and Boundary scenarios.
    • Split into 140 tuning sessions and a reserved 60-session held-out set.
  • Private evaluation sessions

    • Contains 800 organiser-controlled sessions.
    • These sessions were never available to the team.
  • Synthetic test fixtures

    • Small, manually constructed catalogs and fake model responses.
    • Used only to test retrieval, state handling, validation, and failure recovery.

We used no scraped product data, external recommendation datasets, or manually added labels.

Built With

Share this project:

Updates

Submission history