This is a register of the questions about agent selection we have not yet answered. Each entry states why the question matters, the evidence we already have, the evidence we are still missing, and a status: open, being investigated, or answered. It doubles as a preregistration: when the Research Lab sets out to answer one of these, the question and its method are on the record here first. Where a question is already answered, we link the page that answers it.
How large is position bias for each shopping agent, and does it hold in text-only interfaces?
OpenIf early list positions win selection regardless of merit, placement can matter as much as the offer itself, and the effect size should guide where merchants spend effort.
- Evidence we have
- ACES reports that position bias varies by provider and persists even in text-only, 'headless' interfaces, undermining any single notion of a top rank (confirmed in the paper abstract).
- Evidence we're missing
- The exact per-provider regression coefficients from the study's models are not yet verified against the primary source, so we publish no numeric effect size.
Related signals: position bias audit
How much do agent market shares move when a model is updated?
OpenIf a single model update can reorder which products get recommended, a one-time audit goes stale quickly, and merchants need to know how often to re-test.
- Evidence we have
- ACES states that model updates can drastically reshuffle market shares, and documents a model-finalization step that reallocated share and inverted the model's position bias (qualitatively confirmed).
- Evidence we're missing
- The exact numeric magnitude of that reshuffle sits in an appendix we have not yet captured from the primary source, so we publish no figure.
Related signals: retest on model updates
Do query-conditional description rewrites move real agents' choices on live catalogs, not just in simulation?
InvestigatingA large simulated effect is only actionable for merchants if it survives on real storefronts, against production models, on their own catalogs.
- Evidence we have
- ACES reports (v3) that seller-side description rewrites raised average market share substantially across several models in a simulated storefront.
- Evidence we're missing
- A replication on live merchant catalogs measuring the effect on real recommendation share. Our field test is published in methodology and collecting data.
Related signals: description rewrite experiment
Does review depth (count, recency, and star distribution) shift selection beyond a bare average rating?
OpenIf agents weigh how many reviews there are, how recent they are, and how they are distributed, the work is collecting and displaying real review data, not just lifting the average star number.
- Evidence we have
- ACES reports that agents are rating-sensitive and, in simulation, that review count carries its own weight per model (both reported from the paper, not from live agents).
- Evidence we're missing
- A test of whether recency and full star distribution, over and above count and average, change which product a live agent selects. Our review-depth signal is still a labeled inference.
Related signals: review depth, endorsement not sponsored
How differently do live shopping platforms respond to the same price change, and by how much?
OpenIf a markdown that moves share on one engine barely moves another, price strategy has to be tested and budgeted per engine rather than applied as one blended discount.
- Evidence we have
- ACES reports per-model price elasticity in simulation, with Gemini the most price-sensitive and Claude and GPT-4.1 similar (reported, controlled simulation, not real sales).
- Evidence we're missing
- The same measurement against the production agents merchants actually face (ChatGPT, Gemini, Perplexity) on live catalogs, with a per-platform magnitude we can publish.
Related signals: model specific price sensitivity, price competitive awareness
Does publishing an llms.txt file change what any shopping agent does, or is adoption still ahead of measured effect?
OpenMerchants are being told to add llms.txt; they need to know whether it changes discovery or selection or is currently unproven effort.
- Evidence we have
- The llms.txt spec is published and defines a priority order (spec-fact, llmstxt.org). Google states its Search ignores llms.txt entirely (spec-fact), and there is no evidence that shopping agents fetch or act on it.
- Evidence we're missing
- A measurement showing a named shopping agent reads llms.txt and that its presence, or its ordering, changes whether a store is discovered or selected.
Related signals: llms txt priority order, machine readable mirror
Does serving a token-light Markdown mirror of a page measurably improve how often agents cite or select it?
OpenA .md mirror is real engineering; merchants should know whether it moves citation or selection, or is only a plausible best practice for now.
- Evidence we have
- Vendors including Vercel, Cloudflare, and Mintlify report large token savings from serving Markdown to documentation and coding agents (reported vendor examples), and content negotiation is a standard HTTP mechanism (spec-fact).
- Evidence we're missing
- Evidence that any shopping agent requests Markdown, and any measurement tying a .md mirror to higher citation or selection share for a store.
Related signals: machine readable mirror, facts per token
Beyond meeting eligibility, does more complete product structured data measurably shift which product an agent selects?
OpenMerchants routinely pass validation and still lose the pick, so the practical question is whether added depth (shipping, returns, variants, identifiers) changes selection or only unlocks eligibility.
- Evidence we have
- Google documents the fields a listing needs to be eligible and shown (spec-fact), and Adobe reports most retail product pages score only 66 out of 100 for machine readability (reported). Our facts-per-token signal, that fuller completeness then aids selection, is a labeled inference rather than a measured result.
- Evidence we're missing
- A controlled test that isolates structured-data completeness and measures its effect on selection share, held apart from the eligibility it also affects.
Related signals: offer markup completeness, facts per token
Does being ACP- or UCP-ready improve how often an agent selects a store, or only make checkout possible?
OpenProtocol readiness is significant engineering, so merchants need to know whether it wins the pick or is purely table stakes for completing a purchase.
- Evidence we have
- The protocols themselves state that implementing them does not guarantee listings and that each platform runs its own selection process (spec-fact, ACP and UCP docs). Readiness gates whether a purchase can complete, not the comparison that precedes it.
- Evidence we're missing
- A comparison showing whether protocol-ready stores are selected more often than equally-priced peers that are not ready, holding the offer constant.
Related signals: acp readiness, ucp readiness, payment rail readiness
Do live agents penalize a self-applied 'Sponsored' tag and reward genuine endorsements the way the simulation reports?
OpenThe finding is counterintuitive and carries real dark-pattern temptation, so it matters whether it holds on production agents and which legitimate endorsement markers an agent actually recognizes.
- Evidence we have
- ACES reports that agents consistently penalize sponsored tags and reward platform endorsements, with the endorsement lift strongest for Gemini (reported, controlled simulation).
- Evidence we're missing
- Confirmation that live shopping agents show the same penalty and lift, and a mapping of which real-world endorsement markers (editor's pick, verified seller) they read.
Related signals: endorsement not sponsored, review depth