The paper introduces a deterministic, reproducible e‑commerce environment that pre‑commits customer and trajectory parameters, enabling a simulated consumer to attempt purchasing a target cart with the help of an evaluated model. The environment records every assistant action and state, allowing post‑trial evaluation of specific conversation components and applying penalties based on tool‑call accuracy. Using this setup, the authors benchmark eight open‑weight agents (20B–35B parameters) across 160 trials and 44 metrics, revealing nuanced performance issues such as under‑action, over‑purchase, unsupported product attributes, and poor search that are hidden by overall success rates.
By Nimit Shah, Haitz S\'aez de Oc\'ariz Borde
arXiv:2607. 25253v1 Announce Type: new Abstract: Online recommendation has traditionally taken place after a user enters a platform, which determines the candidate pool and the ranking shown to the user.
By Deyao Hong, Kehan Zheng, Qian Li, Jun Zhang, Jie Jiang, Hongning Wang
arXiv:2609.21325v1 Announce Type: new
Abstract: Agentic marketplaces are emerging where AI agents with varying capabilities autonomously complete specialized tasks for buyers. A major challenge of su...
By Steve Drew, Jiayu Zhou
arXiv:2606. 02965v1 Announce Type: new Abstract: Benchmarks for autonomous agents measure whether agents complete tasks, yet this framing is systematically blind to whether an agent should have proceeded at all.
By Victor Ojewale, Suresh Venkatasubramanian
arXiv:2607. 09766v1 Announce Type: new Abstract: AI agents are increasingly deployed in shared environments where they pursue diverse goals and compete for rewards.
By Yaowen Ye, Jacob Steinhardt
ADeptS-Bench is a new benchmark designed to assess the trustworthiness of Computer Use Agents (CUAs) across mobile and desktop devices. It consists of two streams: a Safety stream with paired benign and malicious tasks that embed visual threats, and a Disambiguation stream that tests whether agents seek clarification when instructions are ambiguous. Evaluation of seven models shows none consistently achieves high task success while keeping attack success low, and all models exhibit problematic behaviors such as unhesitant checkout on a $25K order and failure to detect a mislabeled factory reset button.
By Joy Chen, Alejandro Castillejo Munoz, Pierluca D'Oro, Yuxuan Sun, Chloe Evans, Joseph Tighe
arXiv:2606. 30932v1 Announce Type: new Abstract: Two-sided marketplaces connect distinct user groups whose interests often conflict -- improving outcomes on one side could degrade the other side's experience.
By Yufei Wu, Zhen Yan
The paper examines how large language model (LLM) based graphical user interface (GUI) agents respond to digital nudges. Using a randomized online shopping experiment with 3,600 agents across six frontier models, it finds that agents are vulnerable to both automatic and reflective nudges. The study shows that the agents’ reasoning configuration moderates these effects in opposite directions—reducing susceptibility to automatic nudges while increasing it to reflective social influence nudges—and that this redirection is systematically linked to model scale.
arXiv:2608. 11560v1 Announce Type: new Abstract: Personalizing marketing messages with contextual multi-armed bandits (CMABs) drives real business value, yet the objective that ultimately matters - a downstream conversion - is observed only weeks later, too late to drive online learning.
By Sang Su Lee, Vineeth Loganathan, Shishir Dash, Vijay Raghavan
arXiv:2608. 14068v1 Announce Type: cross Abstract: Conversational recommendation for e-commerce is increasingly mediated by large language models (LLMs), yet many real-world deployments operate under a stricter requirement: recommendations must be drawn only from a merchant's fixed catalog, without web search or unsupported product claims.
By Juli Huang, Hannah Clay, Sajjad Beygi, Thomas Sarda, Negin Golrezaei, Amin Saberi
The study examines how large language model (LLM)–based graphical user interface (GUI) agents respond to digital nudges. Using Dual‑Process Theory, researchers tested 3,600 agents across six frontier models in an online shopping experiment and found that the agents were susceptible to both automatic (Type 1) and reflective (Type 2) nudges. The agents’ reasoning configuration moderated these effects in opposite directions: extensive reasoning reduced susceptibility to automatic default nudges but increased susceptibility to reflective social‑influence nudges, with the effect systematically varying by model scale.
By Haya Halimeh, Sascha Kaltenpoth, Kevin B\"osch, Oliver M\"uller
arXiv:2609.05587v1 Announce Type: new
Abstract: Existing evaluations of tool-using agents primarily measure whether an agent can successfully complete diverse tasks with tools. These evaluations gene...
By Hoyeol Yang, Woojung Song, Taewon Kim, Jonghyun Song, Seoyeon Park, Yohan Jo