Inspiration

Product catalogs are usually written for people browsing storefronts—not AI assistants answering nuanced requests about budget, climate, ingredients, values, or exclusions. We created CatalogGym to help merchants understand why their products may be overlooked and improve their data without inventing claims.

What it does

CatalogGym converts catalogs into evidence-linked product passports, tests them against realistic shopper intents, explains recommendation failures, proposes source-grounded repairs, blocks unsupported claims, reruns the same benchmark, and exports agent-ready product data.

How we built it

We built a deterministic Python and FastAPI engine with a responsive JavaScript interface, Supabase/PostgreSQL storage, and Schema.org and OpenAI-compatible exports. Automated Pytest and Node.js tests verify scoring, safety, API contracts, and reproducibility.

Challenges we ran into

One of our biggest challenges was unifying several complex pieces into one reliable workflow: raw CSV ingestion, machine-readable product passports, AI-assisted content generation, evidence-grounded recommendations, Supabase persistence, and simulations across thousands of shopping queries. We also had to preserve tenant isolation, prevent private merchant data or unsupported claims from reaching public agent responses, and provide deterministic fallbacks when an external AI service was unavailable. On the frontend, we worked to make the conversion → comparison → evaluation pipeline understandable without overwhelming users or mixing judge-facing explanations into the operational platform. Finally, because multiple teammates were developing concurrently under a tight deadline, we used isolated branches, regression tests, and careful reconciliation to integrate everyone’s work without overwriting changes.

Accomplishments that we're proud of

CatalogGym delivers the complete visible loop: catalog import, diagnosis, safe repair, blocked-claim review, identical rerun, and structured export. Our reproducible fixture benchmark improved from 31% to 39%, while every promoted claim remained linked to supplied evidence.

What we learned

AI-commerce readiness is not simply better copy. It requires structured data, explicit constraints, traceable evidence, safe claim handling, reproducible evaluation, and clear explanations that both merchants and shoppers can trust.

What's next for CatalogGym

Next, we want to add Shopify and PIM integrations, live source ingestion, managed authentication, monitoring, larger intent suites, and model-assisted extraction—while keeping eligibility rules, provenance, and export safety deterministic and auditable.

Built With

Share this project:

Updates

Submission history