Inspiration

We kept coming back to one asymmetry.

Type "running shoes size 10" into an e-commerce search bar and you get 40,000 results to scroll through. Ask an AI assistant "lightweight shoes for a humid half marathon in Singapore under S$200" and you might get three.

Twenty years of e-commerce content was built around the first case: keywords, titles, specifications, and SEO. Very little was written for the second.

So we asked a simpler question:

What does a product listing need to say for an AI agent to confidently recommend it, and can we actually measure whether it says those things?

That became the foundation of our project.

What We Learned

We expected to find a gap. We did not expect it to be this large.

We audited 300 real Shopify product listings against the information an AI agent needs to reason about a product:

Metric Result
Descriptions mentioning a use case or context 9.8%
Average Content Readiness Score 28.7 / 100
Products scoring above 50 0 / 300

We then looked at the other side.

We generated 2,500 intent-driven shopper queries and ran them through live web search to find products that are currently showing up in AI recommendations. This gave us 474 real products, including brands like NordicTrack, CeraVe, and Rogue Fitness.

Using the same scoring framework, these products averaged just 27.2 / 100, and not a single one scored above 50.

The surprising part was that even products already getting visibility are often winning by accident.

That changed how we thought about the problem. This is not just about helping poorly optimized brands catch up. AI-readable product content is an emerging category and there is still no clear standard for it.

We also found that a complete product record and a useful product record are very different things.

A listing can have a title, category, price, and description and still contain almost none of the context an agent needs to answer questions like:

Who is this for? What situation is it best for? What constraints does it satisfy? When should it not be recommended?

How We Built It

1. A schema as the single source of truth

schema/product_schema.json defines the information a product record should contain.

Both the query generation agents and scoring pipeline use this schema. Adding or changing a field therefore becomes a JSON change instead of a code change.

This lets the same system work across different product categories without rebuilding the pipeline.

2. A deterministic Content Readiness Score

We designed a 100 point Content Readiness Score across five dimensions:

  • Structure: 15
  • Intent Coverage: 35
  • Specificity: 20
  • Negative Space: 10
  • Retrievability: 20

The idea was simple:

If a brand cannot reproduce or understand a score, the score is not very useful.

Where possible, scoring is deterministic. Only the parts that need semantic understanding are sent to a model, and those model-based components are clearly flagged.

Missing measurements do not automatically penalize or inflate a product. The score is normalized over the dimensions that were actually measured:

$$ \text{score} = \left\lfloor 100 \cdot \frac{\sum_{c \in M} s_c} {\sum_{c \in M} w_c} \right\rceil $$

where $M$ is the set of dimensions that were successfully measured.

3. Retrievability as a measurement

We did not want retrievability to be another subjective score.

Our simulation harness generates intent-driven queries for each product category, runs them through the real search pipeline, and records which products appear.

A product's retrievability is simply the fraction of relevant category queries where the product actually appeared.

This gives us a way to measure AI visibility based on observed search behavior instead of an opinion.

4. Benchmarking the products already winning

For competitor analysis, we weight appearances by search rank. Showing up first should count more than showing up fifth.

For rank $r \in {1,\ldots,5}$ and model confidence $s$, across $N$ queries:

$$ w(r) = \frac{6-r}{5} $$

$$ E = \sum_i w(r_i)s_i $$

$$ \text{visibility} = \frac{E}{N} $$

This gives us a benchmark based on observed market visibility rather than simply counting appearances.

5. Category agnostic by construction

The system was designed to work beyond a single vertical.

Running:

python build_benchmark.py --category "skincare"

generates relevant product types and shopper intents for that category, runs them through live web search, and builds a new benchmark database.

A database registry then discovers available datasets automatically, validates their schemas, and exposes their coverage to the interface.

Coverage is discovered, not hardcoded.

Engineering Challenges

The subprocess that worked for exactly one user

Our first implementation spawned a subprocess for every audit request.

It worked fine in a demo. Then two people used it at the same time.

Each audit takes 20 to 40 seconds, so the approach quickly became fragile with concurrent requests. We replaced it with a bounded job queue, added content hash deduplication, and separated request handling from the long running audit.

That was probably the biggest architectural improvement we made.

A database that looked right but wasn't

Two generations of our research data shared table names but had different schemas.

The registry accepted both, but audits would fail halfway through with errors like:

no such column: r.brand

We moved validation to startup. The registry now checks for every column that the audit pipeline actually needs.

A mismatched database gets rejected before it can reach a user request.

Generality that wasn't visible

The system was category agnostic internally, but our implementation initially had a hardcoded path and a directory full of filenames like:

gym_equipment_*

Technically, the system worked for any vertical. Visually, it looked like a treadmill demo.

We replaced the hardcoded assumptions with dynamic dataset discovery, live coverage indicators, and a single command for creating benchmarks across new categories.

We learned that generality has to be visible, not just true.

Never turning competitor evidence into a fact

One of the harder constraints was making sure the system did not accidentally invent product claims.

Competitor evidence can tell us what the market says. It cannot prove that a merchant's product has a particular feature.

So when the audit identifies something potentially useful that has not been verified, it is explicitly marked:

"Folds to [dimensions]"

The audit can suggest a claim, but nothing it suggests can be published as a verified product fact without confirmation from the brand.

Deleting a feature we liked

We initially built a bulk catalog scanner that could process 2,000 rows in 0.15 seconds without making any model calls.

It was fast. It was deterministic. It was accurate.

It also was not very useful.

It mostly told merchants how many rows were missing from each column. We reshaped it twice and eventually cut it.

That taught us an important product lesson:

Precision is not the same as insight.

A feature can be technically impressive and still not help someone make a better decision.

Built With

  • ai-agents
  • css
  • csv
  • e-commerce
  • html
  • javascript
  • json
  • llm
  • multi-agent-systems
  • prompt-engineering
  • pytest
  • rest-api
  • retrieval
  • web-search
Share this project:

Updates

Submission history