# Category and Collection Page Legibility for AI Agents
> Agents often reach products through category and collection pages, so make them server-rendered and crawlable, keep faceted navigation from becoming a crawl trap, and ensure each page truthfully summarizes what it contains.
## Key takeaways
- Agents often reach your products by reading category and collection pages first, so an unreadable listing page can keep otherwise-eligible products out of the comparison.
- Serve the product grid, its links, and the category description in raw server HTML: a JavaScript-assembled listing can reach a non-executing crawler as a near-empty shell.
- Keep faceted navigation legible: give real collections stable URLs, and keep arbitrary filter permutations out of the crawlable set so they don't become a crawl trap.
- A category page is only as useful as it is truthful. It should summarize what the category actually contains, and every product in the grid should belong there.
## Why category pages are an upstream selection lever
Most advice about winning agent selection points at the product page, because that is where the offer lives. But a product page is not usually the first thing an agent reads about your catalog. When someone asks an agent for "waterproof hiking boots under $150," the agent has to assemble a shortlist for the whole category before it ever compares two specific boots, and one of the cheapest ways to assemble that shortlist is to read your collection and listing pages.
We expect category and collection pages to be an upstream selection lever most merchants overlook: an agent building a shortlist for a category query often reaches individual products by crawling and reading listing pages first, so a collection page it cannot parse can keep otherwise-eligible products out of the comparison entirely. This is reasoning from how agents crawl and how the [hub explains agents choose](/how-ai-agents-choose-products/), not a measured effect. But the direction is clear enough to act on: category legibility is the layer that decides whether your products are even in the room when the choosing happens.
This page is about the collection and listing pages themselves. The anatomy of an individual product page is [product detail page anatomy](/product-detail-page-anatomy/); the markup that makes each product parseable is [product schema for AI shopping](/product-schema-for-ai-shopping/); the structured export you hand agents directly is the [product feed](/make-product-feed-ai-readable/). Here the subject is the pages in between.
## Serve collection pages in raw server HTML
The first question about any category page is whether an agent can read it at all. Category pages tend to be the most JavaScript-heavy templates a store runs: infinite scroll, client-side filters, lazy-loaded grids, personalized sorting. Every one of those can move the actual product list out of the initial HTML response.
A collection page whose product grid, category description, and pagination are assembled by client-side JavaScript can render fine to a shopper and still reach a non-executing crawler as a near-empty shell, so the safest default is to serve the grid, its product links, and any intro copy in the raw server HTML. You can check this in seconds: load a top category URL with JavaScript disabled, or view source, and confirm the products and their links are already there. If the grid is empty without JavaScript, an [AI crawler](/glossary/#ai-crawler) that does not execute scripts sees an empty category. The mechanics of which crawlers reach you, and how to keep them unblocked, are covered in [AI crawlers, robots.txt, and llms.txt](/ai-crawlers-robots-llms-txt/).
There is headroom here worth naming. Retail category pages score an average of 74 out of 100 on Adobe's machine-readability index, ahead of product pages at 66 and just behind homepages at 75. Category pages read better than product pages on average, but a score of 74 still means roughly a quarter of a typical listing page is not cleanly machine-readable, and the grid itself is the part most often lost.
One thing you cannot do is route around an illegible category page with a side file. Google's own guidance states that machine-readable text files are not needed to appear in Google Search and that Google Search ignores llms.txt. The page has to carry its own structure. The question of what format to serve, and where content negotiation ends and cloaking begins, belongs to [information density for agents](/handbook/information-density/).
## Structure and name categories the way shoppers ask
Once a category page is readable, its structure is what an agent uses to place it. A flat store where every product hangs off one giant "Shop all" page gives an agent nothing to match a category query against; a clear tree of categories and subcategories gives it a labeled map.
Naming is the same discipline that shapes product titles, applied one level up. We expect a collection to compete better for a category query when its page name and heading use the plain words a shopper would type ("waterproof hiking boots") rather than an internal merchandising label ("The Trailhead Edit"), for the same near-literal query-matching reason that shapes product titles. The evidence behind that matching behavior lives on [product titles that AI agents match](/product-titles-for-ai-agents/); the point here is that a category's own name and URL are query surfaces too, not just decoration for a campaign.
> Illustrative example, not measured data.
A collection built for a seasonal campaign versus one an agent can place:
- **Before:** URL `/collections/the-trailhead-edit`, heading "The Trailhead Edit"
- **After:** URL `/collections/waterproof-hiking-boots`, heading "Waterproof Hiking Boots for Men and Women"
The second version keeps the merchandising story in the intro copy if you want it, but the name, heading, and URL now say what the collection actually is, in the words a query would use.
## Faceted navigation an agent can follow, not a crawl trap
Filters are where category legibility most often breaks. A faceted navigation that generates a fresh crawlable URL for every filter and sort combination can produce thousands of near-identical pages from one real category.
We expect unbounded, crawlable filter permutations (every color, size, and price combination as its own indexable URL) to work against you: they multiply near-duplicate pages that blur which URL represents a real, nameable collection, and they spend a crawler's finite budget on permutations before it has mapped your actual catalog. The fix is not to hide filters from users. It is to decide which filtered views are real collections a shopper would ask for, give those a stable canonical URL, and keep the arbitrary permutations out of the crawlable, indexable set (canonicalizing them to the parent, or excluding them from crawling). The directive mechanics for doing that sit in [AI crawlers, robots.txt, and llms.txt](/ai-crawlers-robots-llms-txt/).
Resist the temptation to solve this by showing crawlers a clean, filter-free version of a page while users get the interactive one. Serving materially different content based on who is asking is cloaking, not optimization, and the boundary between legitimate format negotiation and cloaking is drawn in [information density for agents](/handbook/information-density/).
## Internal links that let an agent map your catalog
Category pages are also the wiring an agent uses to walk your catalog. A plain, crawlable path from category to subcategory to product is what lets it discover every item and understand how they relate.
We expect a product reachable only through on-site search or a JavaScript-driven "load more" button to be effectively invisible to an agent mapping your catalog through category pages, whereas a plain, crawlable link path from category to subcategory to product lets it find and place every item. Two practical consequences follow. First, paginate collections with real, crawlable links rather than a scroll event, so products beyond the first screen are still reachable. Second, hunt for orphaned products: SKUs that no category links to, reachable only via search. The [product feed](/make-product-feed-ai-readable/) is a parallel discovery path that can carry products your navigation misses, but it does not excuse a catalog an agent cannot walk on the site itself. Giving each product one clean, stable, canonical URL keeps that map coherent as filters and campaigns come and go.
## Does your category page truthfully summarize its contents?
The last lever is representativeness. An agent reading a listing page treats it as a summary of what your store carries in that category, and a summary that overreaches reads as a weaker signal.
We expect a category page that misrepresents its contents (an intro claiming the widest range of something over a thin or off-topic grid, or products filed under a category they do not belong to) to read as a less trustworthy summary to an agent using the listing page to judge what your store actually stocks. Keep the category description matched to the grid beneath it: if the copy promises a range, the range should be there. Keep products in categories they genuinely belong to, so a "running shoes" collection is not padded with lifestyle sneakers. The accuracy of each individual product's own claims is a product-page concern, handled through [product schema for AI shopping](/product-schema-for-ai-shopping/) and [product detail page anatomy](/product-detail-page-anatomy/); at the category level, the job is making the page an honest index of what sits under it.
## What to measure
Category legibility is checkable, so check it rather than assume it. On your top category pages, run these:
1. **Raw-HTML render.** Load each top category URL with JavaScript disabled (view source, not devtools) and confirm the product grid, its product links, and the intro copy are all present in the initial response. An empty grid without JavaScript is the finding, not a nuisance.
2. **Orphaned products.** Count SKUs that no category or subcategory page links to, reachable only through search. Those are products an agent walking your catalog may never reach.
3. **Facet URL explosion.** Crawl one busy category and count how many distinct filter and sort permutations are reachable and indexable. Decide which represent real collections worth a canonical URL and which should be kept out of the crawlable set.
4. **Representativeness.** For each top category, confirm every product in the grid genuinely belongs there and the intro copy matches the actual selection, with no promise the grid does not keep.
Track these as baselines you can move and re-measure, the way you would track any ranking signal. How we structure agent-selection measurements, and the metrics we hold ourselves to, is documented in [how we test agent selection](/research/methodology/); real numbers stay in a "data collection in progress" state until the experiments produce them.