# What AI shoppers read before they recommend a product | BYOM blog

URL: https://byom.co/blog/can-ai-shoppers-see-your-store  
Markdown: https://byom.co/blog/can-ai-shoppers-see-your-store.md  
Last updated: 2026-10-02

> Which crawlers visit your store, what each one is for, and the product data, schema and page facts that decide whether an assistant can compare you.

By Kina (Checked by the BYOM team). Published 2026-09-21. 9 minute read. Series: AI shopping.

## Key takeaways

- Blocking an AI company's training crawler is a separate decision from blocking its search crawler, and the published documentation treats them separately.
- Price and availability have to match across the page, the structured data and the feed, or Google disapproves the product.
- Shopify writes basic product markup for you, so check what your theme emits before adding a second copy.

When a shopper asks an AI assistant for a waterproof trail shoe under a set budget, software visits stores on their behalf. It reads what each store publishes, picks the ones it can compare, and leaves the rest out. Most merchants have never looked at who visits, what each visitor is allowed to read, or which of their own pages contradict each other.

This post covers four things you can check: the crawlers that visit and what their operators say they are for, the structured data your Shopify theme writes, the product feed rules Google publishes, and what shoppers themselves look for on a product page.

## Three kinds of visitor, three sets of rules

An AI company that crawls the web usually runs more than one crawler, and each has a different job. The maker of ChatGPT documents three that matter to a shop. GPTBot collects content that may be used in training its models. A separate search crawler surfaces websites in ChatGPT search results. A third fetcher visits a page when a person asks the assistant to open it.

Each answers to robots.txt differently. The documentation says a site that disallows the search crawler will not be shown in ChatGPT search answers, and that it can take about 24 hours after a robots.txt update for its systems to adjust. Disallowing GPTBot, by contrast, indicates that your content should not be used in training. For the fetcher that acts on a person's request, the documentation says robots.txt rules may not apply.

Perplexity publishes the same split. PerplexityBot surfaces and links sites in its search results and, in its own words, is not used to crawl content for AI foundation models. Perplexity User, the fetcher that runs when someone asks a question, generally ignores robots.txt rules because a person requested the visit. Google adds a third pattern. Its Google Extended token controls whether content Google crawls may be used for training future AI models and for grounding in its AI apps, and Google says it does not affect inclusion in Google Search or act as a ranking signal.

| Crawler | What its operator says it is for | What robots.txt does |
| --- | --- | --- |
| GPTBot | Content that may be used in training | Disallow means do not use for training |
| ChatGPT search crawler | Surfacing sites in ChatGPT search | Disallow removes you from those answers |
| ChatGPT User | Visits made when a person asks | Rules may not apply |
| PerplexityBot | Surfacing and linking sites in search | Respects robots.txt |
| Perplexity User | Visits made when a person asks | Generally ignores robots.txt |
| Google Extended | Training and grounding for Google's AI models | No effect on Search inclusion or ranking |

The practical point is that a single blanket Disallow aimed at AI can cost you a place in shopping answers, while a narrow one keeps your content out of training and leaves the search crawlers alone. If you or a developer ever copied an AI blocking snippet into robots.txt, read it against this table.

Googlebot, the crawler behind ordinary Google Search, always respects robots.txt rules for automatic crawls, according to Google's crawler overview. That means the file at the root of your store is the one place where your access rules for the major crawlers live, and it is worth reading once a year, because a rule added during a redesign can sit there unnoticed long after anyone remembers why.

## What Shopify writes for you

Shopify's Liquid language has a structured data filter. Its documentation says the filter converts a product into schema.org markup, as a Product when there is one variant and as a ProductGroup when there are several. It writes JSON LD inside a script tag and includes brand, name, image, offers, price, availability and URL, with identifiers for the product and each offer.

That is a sound base, and it changes the advice. The common mistake is to add a second block of markup from an app or a pasted snippet, which leaves two descriptions of one product that can disagree. Before adding anything, open a product page, view the source, search for ld+json and read what is already there. Whether your theme adds more than the documented fields, such as a GTIN, is something to check page by page rather than assume.

Schema.org defines the Offer type that carries the commercial facts. Its price and priceCurrency properties take an ISO 4217 currency code. Its availability property takes values such as InStock, OutOfStock, PreOrder, SoldOut and LimitedAvailability. The same type has places for a merchant return policy and shipping details, which Google lists among the recommended enhancements for merchant listings, alongside ratings and stock information.

Google separates two classes of markup. Product snippets suit pages where a purchase cannot happen directly, such as an editorial review. Merchant listings suit pages where it can, and they carry the extra detail a shop has: sizing, shipping specifics and return policy. A store page belongs in the second class. Google also says that search result enhancements are shown at the discretion of each experience and may change, so correct markup is a requirement for eligibility and not a promise of appearing.

> **Visible content rule.** Google's structured data policy says not to mark up content that is not visible to readers of the page. A price in the markup that differs from the price on screen is a mismatch a crawler can see.

## The feed rules that apply even if you never run ads

Google Merchant Center publishes the most precise description of what a complete product record looks like, and other systems borrow its vocabulary. Seven attributes are required for every product: an ID, a title of up to 150 characters, a description of up to 5,000 characters, a link, an image link, availability and price. Brand is required for new products, and apparel needs more, including colour and age group in some markets.

Three rules catch small shops out.

### Identifiers

A GTIN is the number under the barcode: 12 digits for a UPC, 13 for an EAN or ISBN, 8 or 13 for a JAN, 14 for a multipack. Google says to submit one only if you are sure it is correct, with a valid check digit, and not to guess or make one up. A product that has no manufacturer GTIN, such as your own brand, should carry a brand and an MPN instead, and the GTIN field stays empty. Google points to the GS1 validation tool for confirming the check digit, so a typo in a spreadsheet does not become a disapproval.

### Price

Outside the United States and Canada, Google says the price should include VAT or GST, and it lists the United Kingdom among the countries where that applies. The amount and currency must match the landing page and the checkout. Shipping goes in shipping settings, not in the product price.

### Availability

The accepted values are in_stock, out_of_stock, preorder and backorder. Google says to match the availability in your data with the landing page and checkout. An in stock product needs an active buy button. Preorder and backorder listings need an expected shipping date. Every variant has to match too, so a size that sold out an hour ago and still shows as available is a disapproval waiting to happen.

A shopping assistant has no reason to be more forgiving than a feed validator. If the same item is in stock in the markup and sold out on the page, one of the two gets distrusted, and you do not choose which.

| Where the fact lives | Price and availability come from | Who reads it |
| --- | --- | --- |
| The product page | What a shopper sees and what checkout charges | People, and any crawler that reads text |
| Structured data in the page | The offers block your theme writes | Search engines and shopping systems |
| A product feed | The feed you submit, such as to Merchant Center | Google shopping surfaces |

Google's own documentation recommends combining structured data on the page with a Merchant Center feed, and notes that some experiences combine data from both when both exist. When they agree, nothing happens. When they disagree, the system has to pick one, and a disapproval or a missing listing is the usual cost. The feed is free to submit for unpaid listings too: Google describes the GTIN as helping make your ad or unpaid listing easier for customers to find.

## What shoppers themselves look for

An assistant reading your page is standing in for a person, and Baymard Institute's 2026 review of product page usability shows where that person gets stuck. The gaps it finds are about facts that a crawler cannot find either.

- **44%** of sites do not display or link to a returns policy on the product page, though 60% of users expect it there (Baymard Institute, 2026)
- **67%** of sites omit a shipping cost estimate on the product page (Baymard Institute, 2026)
- **37%** of sites give no image showing the product in scale, and 42% of users try to judge size from images alone (Baymard Institute, 2026)

Baymard also finds that 57% of sites use dropdown menus for size where buttons would show every option at once, and that 81% fail to show a price per unit for multi quantity products. Each of these is a data question as well as a design one. If the returns window, the delivery cost and the dimensions live only in an image or behind a tab that needs a click, a crawler reading text will not find them. The fix is dull and cheap: put the facts a shopper asks about in the product description as sentences, one fact per sentence, and keep the image for what an image does well. An assistant answering whether a jacket is waterproof can quote a sentence. It cannot quote a photograph.

## What to make of llms.txt

The llms.txt proposal came from Jeremy Howard in September 2024. It asks sites to publish a Markdown file at /llms.txt that gives AI systems a short, curated map: a title, an optional summary in a quotation block, and sections of links with a note on each. The stated reason is that web pages are built for people, with navigation and advertising that make extraction difficult.

It is a proposal, and the crawler documentation cited in this post covers robots.txt rules and does not mention llms.txt. Publishing one costs little and the content is useful to a human too, since it forces you to name the ten pages that explain your shop. Treat it as a tidy summary and do not let it displace the product data work above. The ten pages are usually the same ones a new customer asks about: your best selling collections, sizing, delivery, returns, care and contact.

## A one hour check

- Open yourstore.com/robots.txt and look for any Disallow line aimed at an AI crawler. Decide whether you meant to block training, search, or both.
- View the source of three product pages, search for ld+json, and compare price, currency and availability with what the page shows.
- Pick the product with the most variants and confirm each variant's availability matches its stock.
- Find your GTINs. Check each one has the right number of digits, and leave the field empty where no manufacturer number exists.
- Check that the returns window, return delivery cost and delivery times appear in text on the product page or one click away.

Write down what you find as you go. A short list of mismatches, with the page address beside each, is the useful output of the hour. Fix the price and availability mismatches first, because those are the ones Google says cause disapprovals, then identifiers, then the policy text. Run the same check again after any theme change or app install, since both can rewrite the markup without telling you.

## How BYOM helps

The BYOM Shopify app runs a readiness scan on your installed store through Shopify. It checks catalogue health, shows how AI shoppers see your products, explains why a particular product is invisible to them and reports AI traffic to your store.

Then it proposes the fixes: structured data, FAQ content, llms.txt, pages and redirects. Each one is shown to you first and changes the store only when you confirm it. Every confirmed change is recorded.

The product page: [See the readiness scan](https://byom.co/readiness), [See the Shopify app](https://byom.co/shopify-app).

## Sources

- [Crawler documentation from the maker of ChatGPT, overview of its bots, 2026](https://developers.openai.com/api/docs/bots)
- [Perplexity, guide to its bots and user agents, 2026](https://docs.perplexity.ai/guides/bots)
- [Google Search Central, Google common crawlers, 2026](https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers)
- [Google Search Central, overview of Google crawlers, 2026](https://developers.google.com/search/docs/crawling-indexing/overview-google-crawlers)
- [Google Merchant Center Help, product data specification, 2026](https://support.google.com/merchants/answer/7052112)
- [Google Merchant Center Help, GTIN attribute, 2026](https://support.google.com/merchants/answer/6324461)
- [Google Merchant Center Help, availability attribute, 2026](https://support.google.com/merchants/answer/6324448)
- [schema.org, Offer type](https://schema.org/Offer)
- [Shopify developer documentation, structured data Liquid filter, 2026](https://shopify.dev/docs/api/liquid/filters/structured_data)
- [Google Search Central, introduction to product structured data, 2026](https://developers.google.com/search/docs/appearance/structured-data/product)
- [Google Search Central, structured data general guidelines and policies, 2026](https://developers.google.com/search/docs/appearance/structured-data/sd-policies)
- [Baymard Institute, current state of ecommerce product page UX, 2026](https://baymard.com/blog/current-state-ecommerce-product-page-ux)
- [llmstxt.org, the llms.txt proposal](https://llmstxt.org/)
