arXiv:2510. 12049v4 Announce Type: replace-cross Abstract: We quantify the short-term impact of Generative Artificial Intelligence (GenAI) on sales performance through a series of large-scale randomized field experiments involving millions of users and products at a leading cross-border online retail platform.
By Lu Fang, Zhe Yuan, Kaifu Zhang, Dante Donati, Miklos Sarvary
The paper introduces a hypothesis-driven simulation workflow that screens customer experience (CX) agents before deployment, using synthetic customers and simulated tool outputs to emulate multi-step interactions without accessing production backends. Applied to Nubank’s high-volume Card Delivery and Card Management chat-support agents, the simulation’s binary evaluator scores correlated strongly with production results, and simulation-guided iterations raised transactional net promoter score by 36.69 points in a live A/B test. Additionally, screening over 16,000 simulated conversations helped select a model that increased self‑service rate by 8.82 percentage points without harming net promoter score, demonstrating that simulation enables extensive model exploration safely.
By Edesio Alcoba, Kevin Rossell, Aman Gupta, Shao Tang, Jiwoo Hong, Pabel Carrillo-Mendoza, Wanderson Concei\c{c}\~ao Ferreira, Alvaro Tedeschi, Zayd Simjee, Shreya Rajpal, Bruno Finardi Hime, Christian Sousa, Luis Moneda, Herbert Fei, Daniel Silva, Rohan Ramanath
arXiv:2606. 12924v1 Announce Type: new Abstract: We present a modular two-agent simulation framework for evaluating conversational shopping assistant architectures.
By Jetlir Duraj, Jayanth Yetukuri, Shuang Zhou, Dhruv Varma, Rui Kong, Ishita Khan, Qunzhi Zhou
arXiv:2608. 04205v1 Announce Type: new Abstract: Human evaluation of AI systems and digital products is costly, slow, and difficult to scale.
By Xiaomin Li, Yuexing Hao, Jianheng Hou, Jintao Huang, Qianfeng Wen, Shirley Huang, Yifan Liu, Xiaoyi Liu, Yilan Fan, Yijun Wang, Koutian Wu, Ruoqi Gao, Muhammad Ahmed Mohsin, Jing Tang, Brihi Joshi, Heming Liu, Zheyuan Deng, Zonglin Di, Sankalp Jajee, Jiuyao Lu, Zhiwei Zhang, Saksham Kapoor, Ishan Gupta, Yunhan Zhao, Chanwoo Park, Yucheng Lu, Bing Hu, Weihang Xiao, Aravind Mohan, Hanwen Xing, Runyu Zhang, Mihir Kulshreshtha, Yuanda Xu, Qianyu Zhu, Dianzhuo Wang, Yuxin Xiao, Bowen Jiang, Yongye Su, Wenhao Chai, Zuxin Liu, Lawrence Yunliang Chen, Xuandong Zhao, Ethan Ye, Shivam Patel, Jason Xie, Alex Martin Richmond, Weixiang Ding, Emre Okcular, Diya Mathew, Ziheng Wang, Rana M. Shahroz Khan, Zhejian Peng, Fang Wu, Fan Nie, Xinyang Han, Yubin Kim, Jiawei Zhang, Zhenting Qi, Huangyuan Su, Xu Pan, Abinitha Gourabathina, Hyewon Jeong, Hemanth Neelgund Ramesh, Kumail Alhamoud, Kimia Hamidieh, Zidi Xiong, Samuel Schmidgall, Pengrui Han, Yepeng Huang, Yongheng Wang, Bowen Yang, Alex Gu, Yuchu Wang, Akshay Paruchuri, Brenna Li, Hejie Cui, Jiayuan Ding, Chaosheng Dong, Jiahao Wang, Yixuan He, Chi Wang, Pamela Bhattacharya, Tianyi Peng, Paul Pu Liang, Mitchell Gordon, Yilun Du, Marinka Zitnik, James Zou, Prasanna Tambe, Philip Torr, Emily Fox, Asu Ozdaglar, Dawn Song
The paper describes a production migration of a large-scale customer‑support conversational assistant from a single blended model to a Dynamic Response (DR) architecture. DR replaces the Qwen3‑235B‑A22B responder with a bounded ReAct orchestrator that selects typed tools and a smaller generator that writes from a validated context contract. The migration yields significant improvements: precision‑first entity selection boosts reservation selector precision from 8.3% to 89.1%, typed action IDs eliminate structured‑action hallucination, and hard‑escalation responses drop from 5.60% to 3.08%. Latency is reduced from 3.87 s to 2.24 s, GPU usage is cut by roughly one‑third, and self‑hosting cuts annual model‑serving costs by more than an order of magnitude.
By Cen Mia Zhao, Peng Wang, Chuan Shi, Yufeng Zhang, Ying Lyu, Wanmeng Ren, Robert Xue, Claire Na Cheng, Yashar Mehdad
arXiv:2605.14542v2 Announce Type: replace
Abstract: A skilled live-commerce host is not merely a narrator, but a sales agent who converts viewer curiosity into purchase intent through expert product...
By Yuyan Chen
arXiv:2607. 12085v1 Announce Type: new Abstract: Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, tone, and overall response quality.
By Niranjan Kumar M, Balaji Nagarajan, Karthik Nair, Faysal Satter, Nithin Surendran
arXiv:2609.18729v1 Announce Type: cross
Abstract: Consumers increasingly use AI chatbots for advice on what to buy. With companies like OpenAI and Google monetising their AI through advertising, this...
By Lucas G. Uberti-Bona Marin, Thales Bertaglia, Giovanni Astante, Bram Rijsbosch, Gijs van Dijck, Anik\'o Hann\'ak, Gerasimos Spanakis, Konrad Kollnig
The paper introduces a multi‑agent platform built on CrewAI for conversational business intelligence. Five specialized agents process natural language queries, retrieve and analyze data, generate visualizations via the Model Context Protocol, and deliver actionable insights. The system includes a defense‑in‑depth security architecture, a query parameterization mechanism, and achieves 95.3% functional accuracy with a 24‑second mean latency, outperforming a single‑agent baseline by 22.6 percentage points in accuracy and 20.2% in quality.
By Manoj N M, Vijayakrishna S, Manjunath Srinivas, Rohit Pahan
The paper introduces SalesLLM, a bilingual (Chinese/English) benchmark for evaluating large language models (LLMs) in realistic sales dialogues. It comprises 30,074 scripted configurations and 1,805 curated multi‑turn scenarios from Financial Services and Consumer Goods, with controllable difficulty and personas. An automatic evaluation pipeline uses an LLM judge for sales‑process progress and fine‑tuned BERT classifiers for end‑of‑dialogue buying intent, while a user model, CustomerLM, is trained to improve simulation fidelity. SalesLLM scores correlate strongly with human ratings (Pearson r = 0.86) and reveal that top Chinese LLMs match junior‑to‑intermediate human salespeople but not experts, with cross‑lingual consistency remaining poor.
By Xuanbo Su, Wenhao Hu, Le Zhan, Yuting Xie, Kailin Lyu, Kaijie Chen, Ziwei Li, Yeqiang Wang, Haibo Su, Yunzhang Chen, Ling Huang
The paper introduces a deterministic, reproducible e‑commerce environment that pre‑commits customer and trajectory parameters, enabling a simulated consumer to attempt purchasing a target cart with the help of an evaluated model. The environment records every assistant action and state, allowing post‑trial evaluation of specific conversation components and applying penalties based on tool‑call accuracy. Using this setup, the authors benchmark eight open‑weight agents (20B–35B parameters) across 160 trials and 44 metrics, revealing nuanced performance issues such as under‑action, over‑purchase, unsupported product attributes, and poor search that are hidden by overall success rates.
By Nimit Shah, Haitz S\'aez de Oc\'ariz Borde
The paper introduces a production-ready framework that connects e‑commerce search and CRM systems via AI‑powered Product Research Agents. These agents detect users with exploratory purchase intent, perform multi‑agent research using behavioral data, external knowledge, and catalog information, and then send personalized product recommendations through WhatsApp. In a 23‑day deployment, the system sent about 15,000 notifications, achieving higher click‑through rates than standard campaigns and generating downstream purchases and GMV gains.
By Mandar Kulkarni, Pooja A., Samir Shah