Questions Raised About Reliability of AI Shopping Agents
Researchers at Wharton School have released research findings on the reliability of AI shopping agents. The study confirms that when external information sources such as Wirecutter are referenced, recommended products can change by up to 99 percentage points, and even identical information produces different results when the order of presentation changes. The research reveals that delegating shopping to AI does not necessarily lead to the best purchasing decisions.

Delegating shopping to AI does not guarantee the best choices will be made—researchers at Wharton School (the business school of the University of Pennsylvania) have demonstrated this reality. According to their research, AI shopping agents change their product recommendations significantly when external information sources are even slightly altered or when the order in which information is presented changes.
An AI shopping agent is an AI system that searches for and compares products on behalf of users, proposes purchase candidates, and in some cases automatically performs the purchase operation itself. Platforms including Amazon and Google are actively developing and implementing this functionality, and it is rapidly gaining attention in the e-commerce and retail industries. While the convenience of saving consumers the effort of conducting their own research is appreciated, the reliability question of whether such agents truly operate 'in the consumer's interest' has not been adequately verified to date.
In this research, when information from external sources—specifically from third-party media such as the product review site Wirecutter—was provided, AI agents' product selections were found to change by up to 99 percentage points. Furthermore, even when the information content was identical, changing only the order of presentation resulted in different recommendations. The fact that the 'order' rather than the 'content' of information determines outcomes suggests that the agent's decision-making may not be based on logical and consistent reasoning.
The essence of the problem revealed by these results lies in the 'instability' of AI agents. Humans typically arrive at the same conclusion even when reading the same information in a different order. However, AI agents can have a characteristic tendency toward large fluctuations in output in response to minor changes in input. This is a stability issue called 'robustness,' and current evaluation suggests that caution is warranted when using such systems for purchasing decisions that carry financial consequences.
The implications for consumers are clear. Entrusting product selection and purchasing decisions entirely to an AI agent does not currently guarantee 'obtaining the best judgment.' Particularly when the agent's choice of external information sources and the order in which that information is processed affect outcomes, and when these mechanisms remain opaque to users, the risk of being led toward disadvantageous choices cannot be dismissed.
From an industry and developer perspective, this research poses a fundamental question: how should the 'reliability standards' for AI agents be established? Creating mechanisms to verify whether an agent's recommendations are stable and designing systems that ensure transparency of referenced information sources may become critical tasks in future development. As AI shopping accelerates, how the research's insights are applied to actual product improvements and the formulation of guidelines will be a key point of focus.
This article is an original work independently written and edited by the AI issue editorial team based on factual reporting. © AI issue. Unauthorized reproduction, redistribution, or use for AI training is prohibited.