AI has quickly become part of how consumers research products. Adobe Analytics found that traffic from generative AI tools to U.S. retail sites grew 693.4 percent during the 2025 holiday season compared with the previous year. A year earlier, Adobe found that 39 percent of surveyed U.S. consumers had already used generative AI for online shopping, with product research and recommendations among the most common uses.
The technology companies behind these systems have responded by building shopping directly into their AI products. OpenAI has developed shopping research in ChatGPT to help consumers compare products based on their needs, preferences and budgets. Google uses AI to generate product recommendations and shopping insights, while Amazon has expanded its AI shopping assistant to help consumers research and select products.
As AI takes on a larger role in product discovery, brands have started watching their presence in these recommendations with the intensity they once reserved for Google rankings. Marketers ask AI models the same product questions repeatedly, record which brands appear and track how those results change.
That approach can give marketers a distorted picture of how quickly AI shopping actually moves.
Ask an AI model the same question ten times and you may get ten different combinations of products. A brand can appear first in one answer, fourth in the next and disappear from another. If marketers interpret each answer as a ranking, those results suggest that their position changes constantly.
Generative models introduce variation into their responses, however, and marketers need to account for that variability before deciding that a brand has gained or lost ground. The more useful signal comes from patterns that persist across a large number of relevant shopping conversations.
Data we collected across approximately 3,300 AI shopping recommendation trends involving 186 combinations of brands, prompts and AI models illustrates the difference. Individual answers changed frequently, while the broader distribution of recommendations remained much more consistent.
For marketers, the unit of measurement matters as much as the data itself.
More Queries Can Make AI Shopping Look More Volatile
Search rankings gave marketers a familiar way to measure digital visibility. A result occupied a position at a particular point in time, and marketers could track that position against competitors.
AI recommendations require a different measurement approach because the systems generate answers around a consumer’s query and context. OpenAI, for example, explains that ChatGPT shopping results can take into account the user’s query and contextual information. Google similarly says its shopping results can reflect relevance, search terms, prior activity and shopping preferences.
As a result, a marketer can ask the same general question several times and see different products surface.
Imagine a running shoe brand that appears first in one answer, fourth in another and nowhere in a third. Those three observations make the brand’s performance look unstable. Across hundreds of relevant conversations, however, the same brand might capture roughly the same percentage of recommendations every week.
The distinction becomes important as marketers increase the frequency of their monitoring. Running more queries produces more observations, and more observations create more opportunities to see normal variation. A dashboard can therefore appear increasingly active even when the broader distribution of recommendations changes very little.
Marketers need enough observations to determine whether that movement persists. Without a representative sample, daily monitoring can send them looking for explanations that have little to do with changes in their products, marketing or competitive position.
The consequences extend beyond reporting. Measurement drives decisions about where marketers spend time and money. A company that interprets every daily fluctuation as meaningful may repeatedly change product content, investigate competitors or redirect marketing resources before it has enough evidence to justify those decisions.
Marketers Need to Define What Counts as Meaningful
AI shopping measurement needs a threshold that tells marketers when a change deserves investigation.
A brand disappearing from three answers provides little evidence about its broader performance. If the same brand loses recommendation share across hundreds of relevant shopping conversations and that decline continues for several weeks, marketers have a stronger reason to investigate.
Sample size matters because individual AI responses contain variability. Time matters because persistent patterns provide more information than isolated changes. Marketers should determine both before they use recommendation data to make decisions.
The model also matters.
A brand may gain recommendation share in ChatGPT while its performance in Gemini remains stable. Marketers who combine those observations into a single AI visibility score can lose useful information about where the change occurred.
Different AI shopping systems draw on different information and operate through different product experiences. OpenAI says its shopping research can use merchant product data, publicly available product information and other retail sources. Google says its AI-supported product recommendations draw on Shopping data aggregated from brands, stores and other content providers.
Those differences give marketers a reason to examine recommendation patterns by model rather than assume every AI system will move in the same direction at the same time.
A baseline built over several weeks can provide the reference point. Marketers can then compare new observations with that baseline and investigate changes that persist across a meaningful sample.
Recommendation Share Tells Brands More Than Rank
Recommendation share offers a practical way to measure these patterns. Instead of recording where a product appeared in one response, marketers can measure how often a brand earns a recommendation across a representative set of relevant consumer questions.
Consider a running shoe manufacturer. Its products may appear frequently when consumers ask for shoes for marathon training and rarely when they ask about trail running. That difference tells marketers something useful about the needs that AI systems associate with the brand.
A ranking from an individual answer offers much less context. The difference between appearing second and fifth may disappear the next time someone asks the question. A consistent difference in recommendation share across hundreds of conversations gives marketers a pattern they can study.
Recommendation share can also help marketers identify competitive movement. If a brand typically captures 30 percent of relevant recommendations and begins capturing 40 percent over several weeks, marketers can trace where that increase came from.
They can identify which consumer questions contributed most to the increase and which AI models produced it. From there, they can look at changes in the information available to those systems.
A new product may match a particular consumer need more closely. Retailers may have added more complete product information. Customer reviews or independent coverage may give AI systems additional evidence about when a product fits a particular use case.
The goal of measurement should be to reach that level of analysis. Marketers need metrics that help them identify changes they can investigate rather than metrics that simply produce more movement on a dashboard.
AI Shopping Makes Anecdotes Easy to Mistake for Trends
AI shopping encourages anecdotal measurement because individual answers make compelling evidence.
A screenshot showing a product at the top of a ChatGPT recommendation can circulate through a company within minutes. So can a screenshot showing a major competitor where the company’s own product does not appear. Both examples feel concrete because employees can see the recommendation themselves.
One response still represents one observation.
This problem will grow as consumers rely more heavily on AI during product research. Adobe found that consumers arriving at retail websites from generative AI sources browsed more pages and had lower bounce rates than visitors from other traffic sources. McKinsey has also documented how generative AI assistants can engage consumers earlier in the shopping journey, before they have decided what to purchase.
That growing influence gives brands a strong reason to understand their position in AI recommendations. It also raises the cost of drawing conclusions from weak evidence.
Brands need enough observations to establish how frequently AI systems recommend their products, which consumer questions produce those recommendations and how those patterns develop over time. A baseline gives marketers the context they need to recognize sustained movement.
When marketers identify that movement, they can investigate the information that AI systems use to understand their products. They can examine the accuracy and completeness of product descriptions, retailer information, reviews and independent sources, then determine whether those sources give AI systems enough evidence to connect a product with the needs consumers express.
The growing role of AI in shopping makes this measurement increasingly important. Consumers can now use AI to research products, compare options and narrow a purchase decision before they ever visit a retailer’s website.
Brands need to understand what happens inside those conversations, but more monitoring will not automatically produce better information. The value comes from collecting enough observations to separate ordinary variation from patterns that persist.
Individual AI answers will continue to change. Marketers should expect that. Their attention belongs on the recommendation patterns that persist long enough, and across enough relevant conversations, to justify action.
Credit: Source link



























