Inspiration

Online product content is still designed mainly for humans browsing websites or searching with keywords. However, shopping is increasingly becoming conversational. Instead of searching for “running shoes,” consumers now ask AI assistants questions such as:

“I need lightweight running shoes under S$200 for half-marathon training in Singapore’s humid weather.”

A product may genuinely satisfy this request but still be overlooked because its catalog content does not clearly explain the relevant use case, environment, target customer, or supporting evidence.

This inspired us to build IntentArena—a platform that helps brands understand, test, and improve how AI shopping agents interpret and recommend their products.

What it does

IntentArena turns a product catalog entry into a controlled AI recommendation experiment.

A brand uploads one product text file containing its description, specifications, verified selling points, and claim limitations. IntentArena then:

  1. Creates a Product Digital Twin It extracts structured facts such as category, price, materials, intended users, use cases, performance attributes, and limitations.

  2. Generates relevant Shopper Agents It creates realistic intent-based shopping queries derived from the product category. Each query keeps the uploaded product as a plausible candidate while testing different priorities, such as budget, comfort, distance, climate, or sustainability.

  3. Runs an AI Recommendation Arena The uploaded product competes against relevant mock competitors. The AI scores and ranks the products for one shopper query at a time.

  4. Replays recommendation failures If the uploaded product loses, IntentArena explains why the competing product was easier for the AI to recommend.

  5. Identifies query-specific content gaps Instead of displaying every missing product attribute, the platform ranks only the information required by the current shopper query.

  6. Suggests the minimum content change The brand can reject, edit, or accept a small content improvement without regenerating the entire product description.

  7. Reruns the same query IntentArena keeps the shopper intent and competitor set unchanged, allowing the brand to measure whether the approved content change improved the product’s score and ranking.

  8. Produces a final results report The dashboard compares recommendation wins, average scores, rankings, product versions, and query-level changes before and after optimisation.

Trust Lock

A major concern with generative product content is hallucination. An AI system could make a product appear more attractive by inventing unsupported features or performance claims.

IntentArena addresses this through Trust Lock. Every suggested sentence must be supported by verified information from the uploaded product file. The system may rephrase an approved fact, but it cannot introduce unsupported materials, certifications, medical benefits, performance guarantees, or sustainability claims.

This keeps the brand in control while making every content change traceable to its source.

How we built it

IntentArena uses a full-stack architecture:

  • Next.js and React for the sequential review interface
  • Express and TypeScript for the backend APIs and workflow orchestration
  • OpenAI-powered services for product extraction, shopper generation, competitor evaluation, failure diagnosis, and grounded content suggestions
  • PostgreSQL and Prisma for storing products, shopper agents, simulation results, evidence, and product versions
  • A mock competitor dataset that provides category-relevant products for controlled testing

We separated the AI workflow into specialised stages rather than relying on one large prompt. Product extraction, shopper generation, competition, gap analysis, and content improvement each have their own responsibilities and structured outputs.

Each shopper is processed sequentially, and every accepted or edited suggestion creates a new product version. This allows the final dashboard to compare the original and improved content accurately.

Challenges we faced

One major challenge was ensuring that generated shopper queries were relevant. Completely random queries could make the uploaded product unsuitable from the beginning, producing results that were not useful to the brand. We therefore added a relevance guardrail so that every shopper represents a realistic potential customer for the product category.

Another challenge was maintaining consistency during reruns. To measure content improvement fairly, the same shopper intent and competitor set must be used before and after a change. Otherwise, any score difference could be caused by a different test rather than better content.

We also faced challenges with structured AI output. Product and competitor identifiers had to remain consistent across the model response, database, and frontend. We introduced validation, retries, and controlled identifiers to prevent invalid ranking results.

Finally, we had to balance automation with brand control. Automatically rewriting the entire product page would be fast, but difficult to trust. We instead designed a sequential workflow where the user reviews one grounded suggestion at a time.

What we learned

We learned that AI readiness is not simply about writing longer product descriptions. The most useful content is:

  • Relevant to a specific customer intent
  • Structured for machine reasoning
  • Supported by verifiable evidence
  • Clear about limitations
  • Measurable against competing products

We also learned that recommendation performance should be treated as something brands can test and improve—not as an unpredictable outcome controlled entirely by external AI assistants.

In our recorded prototype run, grounded content improvements increased recommendation wins from 0 out of 3 to 2 out of 3, raised the average product score from 76.7 to 87.3, and improved the average rank from 2.7 to 1.7.

What’s next

The current prototype focuses on running shoes and mock competitors, but the underlying workflow can support other categories through category-specific schemas and evaluation criteria.

Future improvements include:

  • Integrations with Product Information Management systems
  • CSV and API-based bulk catalog uploads
  • Larger competitor datasets
  • Batch simulations across hundreds of shopper intents
  • Category templates for skincare, electronics, fashion, and groceries
  • Multi-model testing across different shopping assistants
  • Exporting approved content changes back into brand catalog systems

Our goal is to help brands measure what AI agents understand, fix only what matters, and prove the improvement—without sacrificing accuracy or control.

Built With

Share this project:

Updates

Submission history