[ research ]

Open Questions in Agent Selection

This open-questions register lists what we do not yet know about agent selection: for each question, the evidence we have, the evidence we are missing, and whether it is open, being investigated, or answered.

This is a register of the questions about agent selection we have not yet answered. Each entry states why the question matters, the evidence we already have, the evidence we are still missing, and a status: open, being investigated, or answered. It doubles as a preregistration: when the Research Lab sets out to answer one of these, the question and its method are on the record here first. Where a question is already answered, we link the page that answers it.

How large is position bias for each shopping agent, and does it hold in text-only interfaces?

Open

If early list positions win selection regardless of merit, placement can matter as much as the offer itself, and the effect size should guide where merchants spend effort.

Evidence we have
ACES reports that position bias varies by provider and persists even in text-only, 'headless' interfaces, undermining any single notion of a top rank (confirmed in the paper abstract).
Evidence we're missing
The exact per-provider regression coefficients from the study's models are not yet verified against the primary source, so we publish no numeric effect size.

Related signals: position bias audit

How much do agent market shares move when a model is updated?

Open

If a single model update can reorder which products get recommended, a one-time audit goes stale quickly, and merchants need to know how often to re-test.

Evidence we have
ACES states that model updates can drastically reshuffle market shares, and documents a model-finalization step that reallocated share and inverted the model's position bias (qualitatively confirmed).
Evidence we're missing
The exact numeric magnitude of that reshuffle sits in an appendix we have not yet captured from the primary source, so we publish no figure.

Related signals: retest on model updates

Do query-conditional description rewrites move real agents' choices on live catalogs, not just in simulation?

Investigating

A large simulated effect is only actionable for merchants if it survives on real storefronts, against production models, on their own catalogs.

Evidence we have
ACES reports (v3) that seller-side description rewrites raised average market share substantially across several models in a simulated storefront.
Evidence we're missing
A replication on live merchant catalogs measuring the effect on real recommendation share. Our field test is published in methodology and collecting data.

Related signals: description rewrite experiment

Does review depth (count, recency, and star distribution) shift selection beyond a bare average rating?

Open

If agents weigh how many reviews there are, how recent they are, and how they are distributed, the work is collecting and displaying real review data, not just lifting the average star number.

Evidence we have
ACES reports that agents are rating-sensitive and, in simulation, that review count carries its own weight per model (both reported from the paper, not from live agents).
Evidence we're missing
A test of whether recency and full star distribution, over and above count and average, change which product a live agent selects. Our review-depth signal is still a labeled inference.

Related signals: review depth, endorsement not sponsored

How differently do live shopping platforms respond to the same price change, and by how much?

Open

If a markdown that moves share on one engine barely moves another, price strategy has to be tested and budgeted per engine rather than applied as one blended discount.

Evidence we have
ACES reports per-model price elasticity in simulation, with Gemini the most price-sensitive and Claude and GPT-4.1 similar (reported, controlled simulation, not real sales).
Evidence we're missing
The same measurement against the production agents merchants actually face (ChatGPT, Gemini, Perplexity) on live catalogs, with a per-platform magnitude we can publish.

Related signals: model specific price sensitivity, price competitive awareness

Does publishing an llms.txt file change what any shopping agent does, or is adoption still ahead of measured effect?

Open

Merchants are being told to add llms.txt; they need to know whether it changes discovery or selection or is currently unproven effort.

Evidence we have
The llms.txt spec is published and defines a priority order (spec-fact, llmstxt.org). Google states its Search ignores llms.txt entirely (spec-fact), and there is no evidence that shopping agents fetch or act on it.
Evidence we're missing
A measurement showing a named shopping agent reads llms.txt and that its presence, or its ordering, changes whether a store is discovered or selected.

Related signals: llms txt priority order, machine readable mirror

Does serving a token-light Markdown mirror of a page measurably improve how often agents cite or select it?

Open

A .md mirror is real engineering; merchants should know whether it moves citation or selection, or is only a plausible best practice for now.

Evidence we have
Vendors including Vercel, Cloudflare, and Mintlify report large token savings from serving Markdown to documentation and coding agents (reported vendor examples), and content negotiation is a standard HTTP mechanism (spec-fact).
Evidence we're missing
Evidence that any shopping agent requests Markdown, and any measurement tying a .md mirror to higher citation or selection share for a store.

Related signals: machine readable mirror, facts per token

Beyond meeting eligibility, does more complete product structured data measurably shift which product an agent selects?

Open

Merchants routinely pass validation and still lose the pick, so the practical question is whether added depth (shipping, returns, variants, identifiers) changes selection or only unlocks eligibility.

Evidence we have
Google documents the fields a listing needs to be eligible and shown (spec-fact), and Adobe reports most retail product pages score only 66 out of 100 for machine readability (reported). Our facts-per-token signal, that fuller completeness then aids selection, is a labeled inference rather than a measured result.
Evidence we're missing
A controlled test that isolates structured-data completeness and measures its effect on selection share, held apart from the eligibility it also affects.

Related signals: offer markup completeness, facts per token

Does being ACP- or UCP-ready improve how often an agent selects a store, or only make checkout possible?

Open

Protocol readiness is significant engineering, so merchants need to know whether it wins the pick or is purely table stakes for completing a purchase.

Evidence we have
The protocols themselves state that implementing them does not guarantee listings and that each platform runs its own selection process (spec-fact, ACP and UCP docs). Readiness gates whether a purchase can complete, not the comparison that precedes it.
Evidence we're missing
A comparison showing whether protocol-ready stores are selected more often than equally-priced peers that are not ready, holding the offer constant.

Related signals: acp readiness, ucp readiness, payment rail readiness

Do live agents penalize a self-applied 'Sponsored' tag and reward genuine endorsements the way the simulation reports?

Open

The finding is counterintuitive and carries real dark-pattern temptation, so it matters whether it holds on production agents and which legitimate endorsement markers an agent actually recognizes.

Evidence we have
ACES reports that agents consistently penalize sponsored tags and reward platform endorsements, with the endorsement lift strongest for Gemini (reported, controlled simulation).
Evidence we're missing
Confirmation that live shopping agents show the same penalty and lift, and a mapping of which real-world endorsement markers (editor's pick, verified seller) they read.

Related signals: endorsement not sponsored, review depth

[ newsletter ]