# An Agent Selection Audit, Worked Through a Fictional Store
> A worked agent-selection audit walks the checklist category by category: for each, you inspect what an agent can read on the store, mark passes and gaps, and fix the highest-leverage gap first. The value is the process, not a score.
> **Illustrative example, not measured data**
>
> Everything below is a walkthrough on an invented store, Thistledown Sleep Co. The store, its products, its prices, and its data are made up to show the process. There are no win rates, no scores, and no measured outcomes on this page, because none have been measured. Any sample title or price is an illustrative field value, not a result.
A real agent-selection audit is a walk through the checklist one category at a time: on each, you inspect what an agent can actually read, mark what passes and what is a gap, then fix the highest-leverage gap before touching anything else. This page shows that walk end to end so you can copy the motion, not a result.
## Key takeaways
- An audit is a process, not a score: inspect what an agent can read, find the gap, fix the highest-leverage one first, then measure whether it moved anything.
- The Agent Selection Checklist has six categories. Working them in order stops you from polishing reviews while a brand-only title quietly loses every match upstream.
- Most gaps are the obvious fields left empty or stale, not exotic ones: a total cost you can only see at checkout, a return window that disagrees with the markup, a title only existing customers would recognize.
- This page diagnoses on a fictional store. Run the same checklist on your own store with the live tool, and hold any fix to a measured agent win rate.
## Meet the store: Thistledown Sleep Co.
> Illustrative example, not measured data.
Thistledown Sleep Co. is an invented direct-to-consumer store selling three lines: a weighted blanket, an organic cotton sleep mask, and a buckwheat pillow. Its snapshot for this walkthrough:
- The blanket ships in weight and size variants labeled "The Nimbus" and "The Nimbus Plus."
- Prices show as plain text, but shipping is "calculated at checkout" and the returns page says "hassle-free returns, conditions apply."
- Product pages carry a star average with no visible count, and descriptions lean on words like "cloud-soft" and "premium."
- No one has ever checked what an AI shopping agent reads on the page, or whether an agent ever recommends the store.
That snapshot is deliberately ordinary. It validates, it converts human shoppers, and it would pass a basic eligibility review. The audit is about the layer eligibility never checks: whether the store wins the [agent selection](/glossary/#agent-selection) once several eligible options sit side by side.
## How to read this walkthrough
Each category below follows the same three-step motion: inspect what an agent can read, separate the passes from the gaps, and name the single highest-leverage fix. The point is the motion. A store that runs this loop honestly beats one that chases a checklist total, because the total hides which gap is actually costing you the pick.
This is a diagnosis, not a re-teaching of every item. Where a fix needs its own chapter, this page hands you to it. To run the same six categories interactively on your own store, with weighted scoring and deep links to each item, use the [Agent Selection Checklist tool](/tools/agent-selection-checklist/). The reasoning behind why these signals move the choice at all lives in [how AI agents choose products](/how-ai-agents-choose-products/).
## 1. Offer legibility
Inspect what an agent can compute before checkout. On a cold product page for The Nimbus, you can read the price as text, so the current-price and parseable-price items pass. But shipping is "calculated at checkout," so the store fails the item that matters most here: total landed cost readable before the agent commits. We expect an agent comparing offers to prefer the one whose all-in cost it can compute up front, and to guess high or drop a merchant whose shipping and fees surface only at the final step.
The variant names are the second gap. "The Nimbus" and "The Nimbus Plus" tell a machine nothing it can map to a shopper's "queen, 15 lb, navy." Unit pricing does not apply to this catalog, so that item is not a gap, just out of scope. The highest-leverage fix is a text line near the price that states total cost to a destination, matching what checkout later charges, followed by renaming variants to lead with their real attributes. The how is in [make your product feed AI-readable](/make-product-feed-ai-readable/), and the checklist item to work is [total landed cost legibility](/tools/agent-selection-checklist/#b-offer-legibility-01).
## 2. Delivery and returns signals
Inspect whether your data answers "when will it arrive?" and "can I return it?" without a human filling gaps. Thistledown states a vague "ships soon" and carries no order cutoff, so an agent cannot derive an estimated delivery date or tell a shopper "order in the next three hours to get it Thursday." Structured shipping data expresses handling and transit as ranges an agent assembles into an arrival estimate; when those fields are vague or absent, we expect the agent to defer to a competitor whose page does answer the date question.
The returns page is worse than empty: its "conditions apply" hedge names no condition, and the prose window does not match the number in the markup. That inconsistency is its own demotion risk, because the agent has to arbitrate between two sources that disagree. The highest-leverage fix is to state concrete handling and transit ranges plus an order cutoff in text, then make the return window, cost, and method agree across prose and structured data. The relevant checklist items are the [order cutoff item](/tools/agent-selection-checklist/#b-delivery-returns-09) and the returns-consistency items beside it.
## 3. Product content for comparison
Inspect completeness against your category, not against the platform minimum. Thistledown filled the required fields and stopped, while the listings it competes against also expose weight, dimensions, fill material, and fabric weight. We expect an agent comparing within a category to favor the item it can actually compare on, so every attribute your peers fill and you leave blank is a comparison you sit out. That is a judgment against peers, which is exactly why no validator flags it.
The descriptions are the second gap. "Cloud-soft" and "premium" answer no concrete buyer question, and they carry no checkable specific an agent can quote. A targeted rewrite can help, but treat it as an experiment, not a guarantee. In the ACES simulation, rewriting a product description shifted recommendation share only in a minority of category-model pairs and did little in most, so gains were concentrated rather than universal. ACES is a controlled simulation over a fixed model set, not live sales. The highest-leverage fix is to fill the attributes category leaders expose and rewrite descriptions to answer the real fit and use-case questions in plain sentences. Work the [category-norms completeness item](/tools/agent-selection-checklist/#b-product-content-01); the title mechanics are in [product titles that AI agents match](/product-titles-for-ai-agents/).
## 4. Trust and policy signals
Inspect whether your credibility reads as depth a machine can weigh. Thistledown shows a bare star average with no count and three testimonials that have not changed in a year. Having a rating is eligibility; the depth behind it is the lever. We expect complete, recent, well-distributed review data to read as more credible depth than a lone average, giving a rating-sensitive agent more to weigh in your favor. A truthful count and recent, dated reviews are the fix, never an inflated figure, which is both a rule violation and a weak signal.
The subtler gap is a leftover "Sponsored" ribbon on some pages from an old promotion, still present in the markup even where shoppers never see it. That framing can actively cost the pick. ACES reports that across its simulated trials, agents consistently penalized a self-applied "Sponsored" tag and rewarded genuine platform endorsements, with endorsement among the strongest content signals it tested. This is a simulation, not measured live behavior. The highest-leverage fix is two-sided: remove the self-applied sponsored styling from organic listings, and pursue real editor's-pick or verified-seller programs you legitimately qualify for, rather than a lookalike badge. Start with the [review-depth item](/tools/agent-selection-checklist/#b-trust-policy-01); the endorsement reasoning is in [how AI agents choose products](/how-ai-agents-choose-products/).
## 5. Price position awareness
Inspect whether you know where your total cost sits, not just your list price. Thistledown sets prices from instinct and has never tracked a competitor's all-in cost for the queries it wants to win. ACES reports price as a real but model-varying lever, so the same markdown can move one engine meaningfully and barely move another. Because it is a simulation over a fixed model set, read the direction, not a fixed magnitude. Without watching competitors' landed prices for your target queries, we expect you cannot judge your position or tell whether a discount is worth its margin.
There is no cloaking shortcut here, and reaching for one fails the audit outright: showing an agent a different price than a shopper sees breaks price-parity rules and is a trust violation, not a growth tactic. The highest-leverage fix is a simple habit, a monthly sheet ranking your landed cost against a few named competitor offers per top query, tested per engine rather than blended. Work the [price-position item](/tools/agent-selection-checklist/#b-price-position-01).
## 6. Measurement
Inspect whether you would even know if a fix worked. Thistledown has never run a shopping prompt against an agent, so every improvement above would ship on faith. That is the gap that makes the whole audit hollow: without measurement, "the fix helped" is a vibe. ACES documents that model updates can drastically reshuffle which products agents choose, and that grid position is a large, provider-specific lever, so a win measured one quarter can quietly evaporate after a release. Both findings come from a simulation, not live sales.
The highest-leverage fix is to define and run the win-rate protocol: a fixed set of realistic prompts, run against each engine on the same day, logging a plain win-or-no-win per engine, with position recorded when results render as a grid. Then re-test after every major model update. Track your [agent win rate](/glossary/agent-win-rate/) as the honest measure of whether any change above earned the pick, using the [win-rate methodology](/research/methodology/), and start from the [win-rate protocol item](/tools/agent-selection-checklist/#b-measurement-01).
## The audit at a glance
The same six-step motion, summarized. Read it as a map of the walkthrough above, not as a verdict on any store.
_Illustrative audit summary for Thistledown Sleep Co. Process, not a score. Every row is a made-up example, not a measured finding._
| Category | What you inspect | What the gap looked like | Highest-leverage fix |
| --- | --- | --- | --- |
| Offer legibility | Can an agent read total cost before checkout? | Shipping only at checkout; variants named 'The Nimbus' | State total landed cost near the price; rename variants to real attributes |
| Delivery and returns | Can it derive an ETA and a return answer? | 'Ships soon', no cutoff; return prose and markup disagree | Concrete handling, transit, and cutoff in text; reconcile return terms |
| Product content | Complete vs category peers, not the minimum? | Missing weight and material; adjectives, no specifics | Fill peer-exposed attributes; answer real buyer questions in prose |
| Trust and policy | Does credibility read as weighable depth? | Bare average, no count; leftover 'Sponsored' markup | Truthful count and recency; remove sponsored styling, earn endorsement |
| Price position | Do you know your landed-cost rank? | Prices set from instinct; no competitor tracking | Monthly landed-cost sheet per top query, tested per engine |
| Measurement | Would you know if a fix worked? | Never measured; no win-rate baseline | Run the win-rate protocol per engine; re-test on model updates |
## Run it on your own store
The walkthrough is the transferable part. Take the same six categories to your own catalog: inspect what an agent reads on a cold page, mark the passes and the gaps honestly, and fix the highest-leverage gap in each before moving on. Do it in order, because a store that polishes reviews while a brand-only title loses every match upstream is optimizing the wrong end.
The [Agent Selection Checklist tool](/tools/agent-selection-checklist/) runs all six categories interactively, with each item's evidence type and deep links to its chapter. When you appear in answers but still lose the pick, the four-cause diagnosis in [shown but not selected](/shown-but-not-selected/) narrows it fast. Whatever you change, hold it to a measured [agent win rate](/glossary/agent-win-rate/) rather than a checklist total, because the total tells you what you touched, and only the win rate tells you whether an agent changed its mind.