Inspiration

People have stopped typing keywords into search bars and started asking agents full questions — "lightweight shoes for a humid half-marathon under S$200." But the product pages those agents read haven't changed at all. They're still written to make a human feel something.

It clicked when we realised the industry had already solved the hard-looking half. OpenAI and Stripe shipped a checkout protocol, Google shipped AP2, Microsoft's NLWeb is turning schema.org into something agents can query. Everyone solved how an agent pays. Nobody solved whether the agent picks you — so good products lose to worse ones purely on how they're described.

What it does

AgentRank scores whether an AI shopping agent could actually recommend your product, rewrites the listing so it can, and then proves the difference instead of asserting it.

Score — six rubric-judged axes: use-case coverage, persona relevance, comparative framing, attribute structure, constraint answerability, trust signal. Not keyword matching — an LLM judging "could an agent act on this and defend the answer?"

Optimize — rewrites from the brand's own facts, plus persona variants. It never fabricates: unknown facts become [weight: confirm], and every inference is surfaced for the brand to verify. The rewrite is re-scored through the identical rubric, so the gain has to survive the same judge.

Ablation — the listing competes against 3–4 strong rivals across 12 shopper queries. Same agent, same rivals, same queries; only the content changes.

Shopper mode — the same failure seen from the shopper's side. A 30-SKU catalogue and a shopper profile, run twice: embeddings shortlist 8, the LLM reranks to 3.

Export — schema.org/Product JSON-LD a brand pastes into their existing product page. No replatform.

How we built it

Next.js 16 App Router on Vercel, AI SDK v7, Zod-validated at every LLM boundary, Recharts, Tailwind v4.

Every stage is a prompt against one content model, all kept in a single lib/prompts.ts. Two model tiers — a reasoning model for the deep single calls, a fast model for the parallel ranking swarm — keep a 36-call ablation inside about 90 seconds, streamed as NDJSON so results fill in live.

The architecture is deliberately category-agnostic. Nothing downstream knows what a running shoe is: category is inferred at scoring time, and the query bank and competitor set for any pasted product are generated at request time. Delete our demo presets and the pipeline still works on anything.

Challenges we ran into

Our first benchmark barely moved, and it was our fault. Raw 1/10, optimized 1/10. The cause: our sample listing had no real specs, so the optimizer correctly refused to invent them — and a rewrite with no facts still can't answer constraint queries. We'd built a product that punished us for our own test data. Rewriting the sample to reflect reality — facts present, buried in prose — is what surfaced the actual insight.

Win-rate was lying to us. It sat flat while mean rank moved meaningfully. On a five-way race, "was it picked #1" throws away most of the signal, so we moved to MRR, Recall@3 and nDCG@3.

We nearly shipped a result we couldn't defend. LLM judges are documented to prefer longer text — and our rewrite is longer. A judge asking "did structure win, or just length?" would have had us.

Avoiding self-marking. If relevance labels are graded from the same content we're optimizing, we're grading our own homework. Every product carries a hidden ground-truth spec sheet the retrieval engine never sees.

Accomplishments that we're proud of

The bloat control came back worse than raw. We padded the original to 5.8× its length with pure marketing prose and zero new facts, to test whether length alone explained our gain.

Arm Picked #1 Recall@3 MRR Words
Raw catalogue 2/12 0.75 0.447 84
Bloat control 0/12 0.67 0.353 491
AgentRank 5/12 1.00 0.694 608

Length doesn't explain the gain — it actively hurts, −21% MRR, while AgentRank gained +55% and made the agent's shortlist on every single query. That one control turns our headline number from a claim into something that survives cross-examination.

We also report a regression: one shopper profile got worse under our own system. A benchmark you always win isn't a benchmark.

What we learned

The gap is framing, not data. Our sample listing already had 238 g, 8 mm drop, 700 km, S$179 — and still scored 48/100. Attributes 76, constraints 72, but persona 18 and comparative 24. Brands aren't missing information; they're missing the structure that makes it reachable. That reframed the whole product: it's a rewrite problem, which means it's automatable.

You have to control for how your judge is broken. Reading the literature on LLM-as-judge — verbosity, position and self-enhancement bias — changed our architecture, not just our confidence. Position bias is why we shuffle candidates with a seeded permutation. Verbosity bias is why the bloat arm exists.

Know what you are. Our instinct was to benchmark against other recommenders. We'd have lost — we don't have a recommender. Holding the recommender constant and varying only the content is both winnable and a stronger claim, because it isolates one variable.

What's next for AgentRank

Real data. The Amazon Shopping Queries (ESCI) dataset — 130k real queries with 2.6M human relevance labels. Swapping LLM-graded ground truth for human ground truth is the biggest credibility upgrade available to us.

Batch processing. Every stage already takes a plain product; CSV upload and catalogue-wide scoring is mostly plumbing. The commercial wedge is a retailer scoring 10,000 SKUs to find which need work first.

Real competitive sets. Rank against a brand's actual competitors from live product pages rather than generated rivals.

The platform play. A pre-onboarding audit for conversational commerce platforms like Rezolve's: before a catalogue goes live, know which SKUs will be invisible.

Built With

Share this project:

Updates

Submission history