AI Shopping Visibility: Tracking What AI Recommends at Product Level
Shopping questions are where AI answers get commercially blunt: ask what to buy and the engines name three products, sometimes with prices, and the conversation often ends there. Brand-level AI visibility doesn't capture this — a brand can be respectfully mentioned while its actual products never make an answer. Tracking has to go to product level, and the technical foundation turns out to be in worse shape than almost anyone assumes. On the real account below, 8.8% of product pages carried product schema. The rest were invisible to any machine trying to quote a price.
By Philipp Enders·Founder, CrunchJunkie·LinkedInBuilds the reporting and AI-visibility tooling this analysis was run with.
The audit that explains a lot of missing AI recommendations: of 34 audited product pages, 3 carried Product schema and none carried review ratings.
Shopping prompts are their own category
"Best LED strip for a workshop", "top robot vacuum under 300 euros", "what should I buy for X" — purchase-intent prompts behave differently from informational ones. The answers are list-shaped, product-named and price-anchored, and the engines lean harder on structured sources: product pages, shopping feeds, review aggregators. A generic prompt set won't surface how you perform here; the shopping set has to be tracked as its own population, with its own metric.
That metric is recommendation rate: of the shopping-prompt runs in which any tracked product was recommended, how often was it yours? It's measured per brand and per product, with the same sampling discipline as everything else — repeated runs, sample sizes shown, and a dash rather than a zero when a product hasn't accumulated enough decisive runs to say anything honest.
Down to the SKU
Brand-level shopping visibility answers "do the engines like us"; product-level answers "which of our products do they actually sell". The gap between the two is routinely the finding. A catalog's hero product can carry the entire recommendation rate while the margin-makers never appear; or engines recommend a product you're phasing out; or they recommend yours — described with a competitor's spec.
Product-level tracking means the catalog lives in the tool: each product with its aliases and SKU variants (engines rarely use your exact product name — matching "V-TAC LED Strip 24V Warm White" requires knowing its aliases), competitor products marked watch-only, and per-product recommendation and win rates over time. The win rate — of runs where any tracked product was recommended, yours won — is the shopping equivalent of share of voice, and the number a merchandising team can actually act on.
The catalog as the engines see it: products with aliases (because AI never uses your exact naming), competitor products on watch, and per-product rates that show dashes until the runs justify a number.
The schema audit most catalogs fail
Here is the supply-side reality check from the account in our hero image. Of 200 product pages, the audit sampled 40; six were unreachable to the checker (its own finding). Of the 34 auditable pages: 3 carried Product schema, 3 carried Offer/price schema, and zero — none — carried review-rating schema. This is a real, trading shop with real revenue.
Why it matters mechanically: when an engine composes "top X under Y euros", price-anchored answers need machine-readable prices. A product page without Offer schema forces the engine to parse prose, and engines under time pressure don't — they quote the competitor whose price was structured. Same for ratings: "best rated" answers are assembled from review markup, and a catalog with zero review-rating schema has opted out of every such answer. The audit turns this from theory into a fix list: which pages, which schema types, in priority order.
Catalog in, answers out
Nobody should type 200 products into a tracking tool. The catalog imports — Google Shopping XML, TSV or CSV feed, or synced directly from Merchant Center — and the prompt side generates from it: shopping prompts built around your actual categories and price points rather than guessed ones.
Worth stating what this doesn't do: importing a feed doesn't make engines recommend you, and we make no such claim. It makes the measurement complete — every product that could be recommended is being watched, so when merchandising asks "does AI ever recommend our mid-range line", the answer is a number with an n, not a shrug.
Where this is heading
Shopping is the AI-visibility surface with a transaction at the end of it, and the engines are building toward that transaction — ChatGPT has shopping results, price comparisons and, increasingly, checkout ambitions; Google's AI surfaces pull straight from Shopping data. When agents buy on behalf of users, recommendation rate stops being a marketing metric and becomes distribution itself.
Which is the argument for measuring now, while your competitors' product pages are as schema-poor as everyone else's: the fixes are unglamorous (markup, feeds, aliases, liftable specs) and the measurement tells you which fix pays. Start with whether the engines can even quote your prices — the free GEO audit includes the schema check — and if you sell more than a handful of SKUs, track the shopping prompts as their own set before the category gets crowded.
Frequently asked questions
Track purchase-intent prompts as their own set and measure at product level, not just brand level: each product with its aliases and SKUs (engines rarely use exact product names), a recommendation rate per product with sample sizes, and competitor products on watch. Catalogs import via Google Shopping feed or Merchant Center so the measurement covers everything that could be recommended.
There's no honest universal benchmark — rates depend entirely on category, prompt set and competition. The usable comparisons are internal: your rate against your tracked competitors on the same prompts, and your trend over time. Any product-level rate shown without its run count should be ignored; too few decisive runs honestly reads as a dash, not a percentage.
Structured data is how engines quote prices and ratings without parsing prose. A page without Offer schema can't be reliably price-quoted in 'best under X' answers; a catalog without review-rating markup is absent from 'best rated' assembly. On a real account we audited, only 8.8% of product pages carried Product schema and none carried review ratings — a fix list, and a common one.
Yes — add competitor products as watch-only entries. They accrue the same recommendation and win rates from the same runs, so 'their mid-range beats our flagship on under-300 prompts' becomes a measured statement with a sample size rather than a hunch from screenshots.
See your AI visibility on your own brand
Reporting and AI search visibility in one console — run your first report and scan inside the 14-day free trial.