Key takeaways
- AI shopping agents rank on structured signals (grid position, price, rating, endorsement, and description text), not on brand recognition.
- The strongest public evidence is the academic ACES framework: a controlled simulation of six frontier models across two 2025 snapshots, choosing from a mock storefront rather than from real sales data.
- Position bias is large but provider-specific and unstable across versions: the grid slot GPT-4.1 favored most is the one GPT-5.1 favors least, and vice versa.
- Rewriting a product description raised share on five of the six models tested but moved Gemini 3.0 Pro Preview essentially not at all, and in a few cases it backfired, so treat rewrites as something to test, not a guarantee.
- The levers you actually control are feed completeness, product schema, titles, and crawler access.
The short answer (signals, not brand habit)
When an AI shopping agent picks a product, it is reading a scorecard, not recalling a brand. In the strongest evidence available today (the academic ACES framework, covered below), the levers that moved an agent's choice were all structured attributes: where the product sat in the results grid, its price, its star rating, whether it carried a platform endorsement, and how its description was written. That is why the question of agent selection (why an agent chooses one store or offer over another) comes down to legible product data, not reputation.
We read the evidence to mean that an agent's pick is driven by machine-legible product signals rather than the brand familiarity that sways human shoppersHypothesis (our analysis), which is how a small store with complete, well-structured data can beat a household name it would never outrank in a person's memory. This page is the selection chapter of the broader agentic commerce optimization guide; the platform playbooks then show how each engine weights these signals in practice.
What the evidence shows: the ACES framework
The most rigorous public work on how buying agents choose is a 2025 study whose framework is called ACES. ACES, introduced in the paper "What Is Your AI Agent Buying?", was authored by Allouah and colleagues across MyCustomAI, Columbia Business School, and Yale, and is a provider-agnostic method for auditing agent decisions: a mock storefront is shown to vision-language-model shopping agents in repeated randomized trials, and their selections are analyzed with regressionSpec-factACES, Allouah et al., arXiv:2508.02630. Because the storefront and every attribute are controlled, the study can isolate what actually shifts a choice. The work has also cleared peer review: a short-paper version of ACES was published in the Proceedings of the ACM Web Conference 2026 (WWW '26) on 12 April 2026Spec-factACM Digital Library, Proceedings of the ACM Web Conference 2026; the figures on this page continue to cite the fuller arXiv revision, which carries the complete tables.
Keep two caveats in mind before the numbers. The paper's December 2025 revision evaluates six focal models at two snapshots: an August 2025 cohort (Claude Sonnet 4, GPT-4.1, and Gemini 2.5 Flash) and a December 2025 cohort of their successors (Claude Opus 4.5, GPT-5.1, and Gemini 3.0 Pro Preview), all inside a controlled simulation rather than a record of real-world salesReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). And every effect it found is model-specific: the same signal has a different weight on each agent, and, as the newer cohort shows, on each generation. The sections that follow report the paper's numbers exactly, each one carrying that simulation caveat.
Position & layout bias
The single largest effect ACES found is where a product sits on screen. For a Claude Sonnet 4 agent, a product in the bottom-right corner of the grid was selected about 4.5% of the time; moving that same product to the second or third column of the top row raised its selection rate roughly fivefold, while moving it to the top-left corner produced only about half that increaseReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). This is two-dimensional grid position, not a one-dimensional list rank.
And the effect is not universal. Position bias varied by provider and persisted even in text-only, "headless" interfaces, undermining any single notion of a "top" rankReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). For a merchant that means placement inside an agent's result surface is powerful but rarely something you set directly; what you control is being eligible to appear at all, which loops back to feed quality and crawler access. See position bias in the glossary.
The newest evidence shows that bias can flip outright between versions. Comparing the two cohorts, GPT-4.1's top-row coefficient was +1.045 (Table 2), but its successor GPT-5.1's is −0.701 (Table 3): the paper describes the two generations' position biases as almost opposite, with the slot GPT-4.1 favors most being the one GPT-5.1 favors least and vice versa, while Gemini 3.0 Pro Preview's top-row coefficient is strongly positive at 2.146ReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). A model update did not merely resize the position bias here; it reversed its sign.
Price sensitivity
Agents respond to price, and how strongly depends on the model. For the August 2025 cohort, ACES estimated log-price (ln Price) coefficients of −1.623 for Claude Sonnet 4, −1.612 for GPT-4.1, and −2.190 for Gemini 2.5 Flash, negative everywhere (cheaper wins, all else equal), with Gemini the most price-sensitive of that cohortReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). In the December 2025 cohort the ln Price coefficients are −1.886 for Claude Opus 4.5, −2.798 for GPT-5.1, and −2.249 for Gemini 3.0 Pro Preview, all negative, making GPT-5.1 the most price-sensitive of all six modelsReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). Price competitiveness matters across agents, then, but the same discount buys more selection lift on some engines than others. Because these coefficients are model-specific and drawn from a simulation, treat them as directional evidence that price is a real lever, not as a pricing formula.
Ratings & review depth
Star ratings move selection, and they move it most on the most rating-sensitive models. For the August 2025 cohort, raising a product's rating by 0.1 stars lifted a baseline 10% selection probability to 15.4% for Claude Sonnet 4, 20.3% for GPT-4.1, and 16.0% for Gemini 2.5 Flash, a large swing for a small rating change, with GPT-4.1 the most rating-sensitiveReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17).
ACES also models review count as a signal distinct from the average score, and reports it as a positive selection factor across the models testedReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17), which points to review depth mattering beyond the headline rating. In the paper's conditional-logit estimates (Table 2 of the current revision), the log review-count coefficient is positive for every model tested: 0.415 for Claude Sonnet 4, 0.739 for GPT-4.1, and 0.501 for Gemini 2.5 Flash, each statistically significant (p below 0.001), with GPT-4.1 weighting review depth most heavilyReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). The December 2025 cohort weights both signals in the same direction: rating coefficients of 11.148 for Claude Opus 4.5, 9.247 for GPT-5.1, and 4.224 for Gemini 3.0 Pro Preview, and log review-count coefficients of 0.979, 0.800, and 0.669 respectively, all statistically significant (p below 0.001)ReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). The paper reports these newer figures only as coefficients, so there is no baseline-lift percentage to quote for the December models. As with the price coefficients above, these are model-specific logit coefficients from a simulation: directional evidence that more reviews raise selection odds, not a percentage lift you can bank. Either way, genuine ratings and review volume remain among the few selection levers you can influence honestly and directly.
The "Sponsored" penalty & endorsement lift
Two counterintuitive findings sit together: agents penalize paid-looking tags and reward platform endorsements. For the August 2025 cohort, tagging a product "Sponsored" pushed a baseline 10% selection probability down to 8.9% (Claude Sonnet 4), 8.0% (GPT-4.1), and 7.9% (Gemini 2.5 Flash), a consistent penalty across all three modelsReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). In the other direction, the same cohort's "Overall Pick" endorsement lifted the baseline 10% to 24.3% (Claude), 19.9% (GPT-4.1), and 42.6% (Gemini), with Gemini rewarding endorsement most stronglyReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17).
Both effects survive the generation change and, if anything, grow stronger. In the December 2025 cohort the Sponsored coefficients are −0.340 (Claude Opus 4.5), −0.371 (GPT-5.1), and −0.623 (Gemini 3.0 Pro Preview), each more negative than its August predecessor, while the Overall Pick coefficients rise to 1.865, 1.342, and 2.144 respectivelyReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). The scarcity tag ("Only X Remaining") was negative or null for the August 2025 cohort, significantly so only for Gemini 2.5 Flash (−0.342), but for Gemini 3.0 Pro Preview it flips to slightly positive and significant, at 0.219 (p below 0.05)ReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). This is evidence that a model generation can flip even the sign of an effect.
The honest implication (and it is an inference, not a growth hack) is that the winning move is to earn legitimate endorsement signals, not to fabricate them; manufacturing badges or gaming "pick" markers is exactly the dark pattern this result would tempt a merchant intoHypothesis (our analysis).
The description-rewrite experiment
The one lever on this list a merchant fully controls is the words on the product, and ACES tested it directly. A minimal AI seller agent rewrote a single randomly chosen focal product's description in each category, using competitor sales data as input; GPT-4.1 acted as the seller agent optimizing against the August 2025 buyer models, and GPT-5.1 against the December 2025 buyer modelsReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17).
The rewrite moved share on most, but not all, of the buyers. Averaged across categories, the focal product's market share rose by +3.66 percentage points on Claude Sonnet 4, +8.37 pp on GPT-4.1, +14.79 pp on Gemini 2.5 Flash, +7.38 pp on Claude Opus 4.5, and +14.89 pp on GPT-5.1, all statistically significant; on Gemini 3.0 Pro Preview it moved +0.32 pp, which was not statistically significant, leaving the effect significant in five of the six buyer modelsReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). On the newest Gemini, the one-shot rewrite did essentially nothingReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). That single null result is enough to reject any pitch that description rewrites reliably work everywhere: on at least one current frontier model, this one did not.Hypothesis (our analysis)
The averages also hide how uneven the effect is. Across category-model pairs spanning all six buyer models, 67% saw no statistically significant change from the one-shot rewrite, while the remaining 33% produced large gainsReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). One category shows what a large gain looks like. An office lamp was the only category with consistent, significant gains across all six buyers, ranging from +7.1 pp on Gemini 3.0 Pro Preview to +80.4 pp on GPT-5.1ReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). The authors hypothesize the mechanism: the original description did not place the query keyword "Office" early, so it was truncated on the mock storefront and the agents overlooked it; the rewrite front-loaded "Office"ReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). That is the same reason where a keyword sits in your title and feed text can decide whether an agent ever registers it.
Rewrites are not risk-free. In a few cases the AI-rewritten description reduced market share, for a stapler under Claude Opus 4.5 and a mousepad under Gemini 3.0 Pro PreviewReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). The lesson is continuous testing, not a set-and-forget rewrite.
Behavior is model-dependent (and changes on updates)
Every effect above has a different size on each model, and the ranking is not stable over time. ACES found agents concentrate demand on a handful of "modal" products and never select some brands at all, with the pattern differing by modelReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). More unsettling for anyone optimizing to a single engine: the authors documented that a model update can drastically reshuffle market shares, as it did between Gemini 2.5 Flash Preview and Gemini 2.5 FlashReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17).
The December cohort makes that reshuffle concrete. In the fitness-watch category, the Fitbit Inspire's share of agent picks rose from 45% under Claude Sonnet 4 to 77% under Claude Opus 4.5, while falling from about 25% under GPT-4.1 to 6% under GPT-5.1. Among iPhone 16 Pro covers, GPT-4.1's modal pick (Mikeke, 62.6%) collapsed to 5% under GPT-5.1, which instead picks ESR 95% of the timeReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). The paper also observes that this choice homogeneity, with agents piling onto a few products, is more pronounced in the latest modelsReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17).
This is why AgentMint.net types every model-behavior claim by source and date: what an agent rewards this quarter it may weight differently after the next release. Treat agent behavior as a moving target: the per-engine platform playbooks track where the engines currently diverge.
Baseline rationality has matured
The newer models are better at the basics. On single-attribute tests, the share of trials where an agent failed to detect a +0.1-star advantage fell from 28.7% for Claude Sonnet 4 and 15.1% for GPT-4.1 in August 2025 to below 1.7% for the December 2025 models, which the paper reports almost never miss such single-dimension comparisonsReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). The reading for merchants: agents now reliably read the scorecard, so getting scored correctly is table stakes; what stays model-dependent and durable is how each agent weights that scorecard (position, badges, and the attribute sensitivities above), so legibility gets you scored and weighting decides who wins.Hypothesis (our analysis)
What this means for merchants: the practical levers
You cannot set your grid position or force an endorsement, but you can control the inputs those signals are computed from. The levers below translate the evidence above into work you can start this week:
- Make your product feed AI-readable: completeness and structure are what make you eligible to be scored at all.
- Ship product schema (JSON-LD) for AI shopping: the machine-readable layer agents trust when your page and feed disagree.
- Write product titles AI agents match: the description-rewrite experiment above showed title and description text can move share by double digits on some models and next to nothing on others, so the words are a real but uneven lever.
- Manage AI crawlers, robots.txt and llms.txt: block the crawlers and you delete the supplemental page data agents read beyond your feed.
- Deliver those signals fast and lean: even complete data loses if an agent cannot fetch it quickly or has to dig through filler to reach it. Two infrastructure chapters cover that layer, serving agent traffic (caching and reachability, so being open to agents does not cost you) and raising information density (packing the decision facts into fewer tokens).
Together these are the inputs to the scorecard. The platform playbooks then show how ChatGPT, Gemini, and Perplexity weight them differently.
Caveat: simulation, not sales
Every number on this page comes from one controlled study, and its limits matter as much as its findings. ACES is a simulation: vision-language-model agents choosing from a mock storefront in randomized trials, across a fixed set of six models at two snapshots (Claude Sonnet 4, GPT-4.1, and Gemini 2.5 Flash in August 2025; Claude Opus 4.5, GPT-5.1, and Gemini 3.0 Pro Preview in December 2025), not a record of real purchases on live storesReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). Read these effects as evidence of what kinds of signals agents weigh and how differently they weigh them, not as guaranteed conversion math for your catalog.
One methodological detail bounds how far the numbers travel. ACES ran every agent at minimal or disabled reasoning (thinking off for Claude Sonnet 4 and Gemini 2.5 Flash, a 500-token thinking budget for Claude Opus 4.5, and low reasoning effort for GPT-5.1 and Gemini 3.0 Pro Preview) at temperature 1.0, settings chosen to mirror latency-constrained deploymentsReportedACES, Allouah et al., arXiv:2508.02630 (2025-12-17). Production assistants that reason for longer or search the web may weight these signals differently, so read the coefficients as a minimal-reasoning baseline rather than a description of every chat-based shopping assistant.Hypothesis (our analysis)
ACES is also no longer the only testbed pointing this way. ABxLab, a framework from a team including MIT Media Lab researchers, independently probes agent purchasing choices by manipulating price, ratings, and persuasive nudges under controlled conditions, and reports that agent decisions shift predictably and substantially in response.ReportedABxLab, Cherep et al., arXiv:2509.25609 (2026-02-24) A second simulation reaching the same broad conclusion (structured signals, deliberately varied, move what an agent buys) does not remove the simulation caveat, but it does mean the pattern no longer rests on a single study.
The methodology is inspectable (the ACES simulator is released under the MIT licenseSpec-factACES simulator, github.com/mycustomai/ACES), and AgentMint.net's own research lab is replicating the description-rewrite finding on live catalogs, with results held in a "data collection in progress" state until real numbers exist. Until then, treat ACES as the best available map of how agents choose, and your own measured agent win rate as the territory.