Inspiration
AI-powered search is changing how people discover products, but most online listings are still written for traditional search engines and marketplace filters. We wanted to answer a practical question: what information actually helps an AI system find and recommend a product?
We can't honestly claim to predict rankings on live platforms, because those systems are closed and change without notice. So we built a controlled testbed instead: different versions of the same listing compete under identical conditions, and every run can be repeated exactly.
What it does
The GEO Lab takes a real product listing and its verified product facts, then builds several truthful variants of it:
- A clearer, better-structured listing using only facts already present
- A complete specification listing exposing every supplied fact
- A benefit-focused listing
- A version aligned with the language shoppers search in
- Single-fact variants that add back one missing detail at a time
Two controls run alongside them. The identity control is a byte-identical copy of the original, and the run fails if it ever retrieves differently, which tells us the harness itself is not adding noise. The misleading control inflates one numeric fact on purpose, so our claim validator has something it is required to catch.
Every version faces the same benchmark peers and the same shopping queries. The report answers four questions: what happened, why, what it means, and what the brand should do.
How we built it
Python, FastAPI, SQLite, NumPy, and any OpenAI-compatible structured-output endpoint, with Gemini as our configured default.
Retrieval fuses three deterministic channels with reciprocal rank fusion: SQLite FTS5 BM25, a pinned semantic projection, and typed-attribute matching, with canonical hard constraints applied on top. The model then judges recommendation quality after seeing the same candidate set. We report these separately, because a listing can be easy to find without being persuasive, and a single blended score would hide exactly the disagreement we care about.
Runs are staged to keep evidence honest and API usage bounded:
- Screen. Up to 10 variants across 6 generated queries, capped at 60 recommendation calls.
- Confirmation. Only the best two valid variants advance, retested on 12 held-out queries against the original and the identity control.
Bias controls are built into the harness. The candidate set is frozen per query, our product rotates through every position across queries and trials, condition order alternates, and temperature is fixed at zero. Query generation, variant construction, retrieval, metrics, and promotion are all deterministic. The model is used only for recommendation judgements and the final explanation, and its rewrite is rejected if it contains any number that cannot be traced to a supplied fact.
Challenges
Separating visibility from recommendation quality was the hardest part. Early results looked contradictory: a variant could rank higher in retrieval and still lose once the model compared it against peers. Those results don't conflict. They point to different content gaps, and collapsing them into one number would have destroyed the finding.
We also had to bound API usage under low quota, handle rate limits with resumable and cancellable runs, and keep the report readable without hiding the metrics behind it.
What we learned
Better keyword alignment can make a product easier to find, but that alone does not make it more recommendable. Getting chosen also takes clear specifications, relevant benefits, and a legible reason to prefer this product over the alternatives.
GEO is about making verified product information easy for AI systems to retrieve, understand, compare, and confidently recommend. Keyword insertion is a small part of that.
Limitations
We are deliberate about what this does not show. Results are directional and descriptive, not confirmatory, until sample sizes and decision rules are preregistered. The benchmark peers are deterministically generated rather than scraped from a live marketplace. Queries are templated from the product's own facts, so we need independently written queries before we can rule out favorable framing. The semantic channel currently uses a pinned deterministic projection rather than a trained sentence encoder, which preserves reproducibility but is our main remaining deviation from the intended design. And nothing here measures public ChatGPT visibility, indexing, conversions, or sales.
What's next
Larger catalogues, more product categories, additional models, and independently written queries, so we can find out which of these effects survive outside the current benchmark.

Log in or sign up for Devpost to join the conversation.