arXiv AI

Generative AI in Action: Field Experimental Evidence from Alibaba's Customer Service Operations

arXiv:2603. 29888v2 Announce Type: replace-cross Abstract: In collaboration with Alibaba, we study how a generative AI assistant affects service performance in e-commerce after-sales operations.

arXiv AI
Jun 2

Generative AI and Sales Productivity: Field Experiments in Online Retail

arXiv:2510. 12049v4 Announce Type: replace-cross Abstract: We quantify the short-term impact of Generative Artificial Intelligence (GenAI) on sales performance through a series of large-scale randomized field experiments involving millions of users and products at a leading cross-border online retail platform.

By Lu Fang, Zhe Yuan, Kaifu Zhang, Dante Donati, Miklos Sarvary
arXiv AI
Sep 25

Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale

The paper introduces a hypothesis-driven simulation workflow that screens customer experience (CX) agents before deployment, using synthetic customers and simulated tool outputs to emulate multi-step interactions without accessing production backends. Applied to Nubank’s high-volume Card Delivery and Card Management chat-support agents, the simulation’s binary evaluator scores correlated strongly with production results, and simulation-guided iterations raised transactional net promoter score by 36.69 points in a live A/B test. Additionally, screening over 16,000 simulated conversations helped select a model that increased self‑service rate by 8.82 percentage points without harming net promoter score, demonstrating that simulation enables extensive model exploration safely.

By Edesio Alcoba, Kevin Rossell, Aman Gupta, Shao Tang, Jiwoo Hong, Pabel Carrillo-Mendoza, Wanderson Concei\c{c}\~ao Ferreira, Alvaro Tedeschi, Zayd Simjee, Shreya Rajpal, Bruno Finardi Hime, Christian Sousa, Luis Moneda, Herbert Fei, Daniel Silva, Rohan Ramanath
arXiv AI
Aug 6

MatrAIx: Simulating the World with 8.3 Billion Persona Agents

arXiv:2608. 04205v1 Announce Type: new Abstract: Human evaluation of AI systems and digital products is costly, slow, and difficult to scale.

By Xiaomin Li, Yuexing Hao, Jianheng Hou, Jintao Huang, Qianfeng Wen, Shirley Huang, Yifan Liu, Xiaoyi Liu, Yilan Fan, Yijun Wang, Koutian Wu, Ruoqi Gao, Muhammad Ahmed Mohsin, Jing Tang, Brihi Joshi, Heming Liu, Zheyuan Deng, Zonglin Di, Sankalp Jajee, Jiuyao Lu, Zhiwei Zhang, Saksham Kapoor, Ishan Gupta, Yunhan Zhao, Chanwoo Park, Yucheng Lu, Bing Hu, Weihang Xiao, Aravind Mohan, Hanwen Xing, Runyu Zhang, Mihir Kulshreshtha, Yuanda Xu, Qianyu Zhu, Dianzhuo Wang, Yuxin Xiao, Bowen Jiang, Yongye Su, Wenhao Chai, Zuxin Liu, Lawrence Yunliang Chen, Xuandong Zhao, Ethan Ye, Shivam Patel, Jason Xie, Alex Martin Richmond, Weixiang Ding, Emre Okcular, Diya Mathew, Ziheng Wang, Rana M. Shahroz Khan, Zhejian Peng, Fang Wu, Fan Nie, Xinyang Han, Yubin Kim, Jiawei Zhang, Zhenting Qi, Huangyuan Su, Xu Pan, Abinitha Gourabathina, Hyewon Jeong, Hemanth Neelgund Ramesh, Kumail Alhamoud, Kimia Hamidieh, Zidi Xiong, Samuel Schmidgall, Pengrui Han, Yepeng Huang, Yongheng Wang, Bowen Yang, Alex Gu, Yuchu Wang, Akshay Paruchuri, Brenna Li, Hejie Cui, Jiayuan Ding, Chaosheng Dong, Jiahao Wang, Yixuan He, Chi Wang, Pamela Bhattacharya, Tianyi Peng, Paul Pu Liang, Mitchell Gordon, Yilun Du, Marinka Zitnik, James Zou, Prasanna Tambe, Philip Torr, Emily Fox, Asu Ozdaglar, Dawn Song
arXiv AI
Sep 10

From Monolithic Blending to Agentic Orchestration: Dynamic Response for Conversational Assistants at Scale

The paper describes a production migration of a large-scale customer‑support conversational assistant from a single blended model to a Dynamic Response (DR) architecture. DR replaces the Qwen3‑235B‑A22B responder with a bounded ReAct orchestrator that selects typed tools and a smaller generator that writes from a validated context contract. The migration yields significant improvements: precision‑first entity selection boosts reservation selector precision from 8.3% to 89.1%, typed action IDs eliminate structured‑action hallucination, and hard‑escalation responses drop from 5.60% to 3.08%. Latency is reduced from 3.87 s to 2.24 s, GPU usage is cut by roughly one‑third, and self‑hosting cuts annual model‑serving costs by more than an order of magnitude.

By Cen Mia Zhao, Peng Wang, Chuan Shi, Yufeng Zhang, Ying Lyu, Wanmeng Ren, Robert Xue, Claire Na Cheng, Yashar Mehdad
arXiv AI
Jul 15

Operationalising Multi-Dimensional Evaluation for Conversational Agents: A Scalable, Governed Pipeline with Selective Re-evaluation and Model Benchmarking

arXiv:2607. 12085v1 Announce Type: new Abstract: Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, tone, and overall response quality.

By Niranjan Kumar M, Balaji Nagarajan, Karthik Nair, Faysal Satter, Nithin Surendran
arXiv Computation and Language
Sep 17

"If I Had to Buy Just ONE: Galaxy S26 Ultra": Auditing AI-Generated Product Recommendations

arXiv:2609.18729v1 Announce Type: cross Abstract: Consumers increasingly use AI chatbots for advice on what to buy. With companies like OpenAI and Google monetising their AI through advertising, this...

By Lucas G. Uberti-Bona Marin, Thales Bertaglia, Giovanni Astante, Bram Rijsbosch, Gijs van Dijck, Anik\'o Hann\'ak, Gerasimos Spanakis, Konrad Kollnig
arXiv AI
Aug 20

A Multi-Agent Platform for Automated Enterprise Analytics and Insight Generation

The paper introduces a multi‑agent platform built on CrewAI for conversational business intelligence. Five specialized agents process natural language queries, retrieve and analyze data, generate visualizations via the Model Context Protocol, and deliver actionable insights. The system includes a defense‑in‑depth security architecture, a query parameterization mechanism, and achieves 95.3% functional accuracy with a 24‑second mean latency, outperforming a single‑agent baseline by 22.6 percentage points in accuracy and 20.2% in quality.

By Manoj N M, Vijayakrishna S, Manjunath Srinivas, Rohit Pahan
arXiv Computation and Language
Aug 27

Sell More, Play Less: Benchmarking LLM Realistic Selling Skill

The paper introduces SalesLLM, a bilingual (Chinese/English) benchmark for evaluating large language models (LLMs) in realistic sales dialogues. It comprises 30,074 scripted configurations and 1,805 curated multi‑turn scenarios from Financial Services and Consumer Goods, with controllable difficulty and personas. An automatic evaluation pipeline uses an LLM judge for sales‑process progress and fine‑tuned BERT classifiers for end‑of‑dialogue buying intent, while a user model, CustomerLM, is trained to improve simulation fidelity. SalesLLM scores correlate strongly with human ratings (Pearson r = 0.86) and reveal that top Chinese LLMs match junior‑to‑intermediate human salespeople but not experts, with cross‑lingual consistency remaining poor.

By Xuanbo Su, Wenhao Hu, Le Zhan, Yuting Xie, Kailin Lyu, Kaijie Chen, Ziwei Li, Yeqiang Wang, Haibo Su, Yunzhang Chen, Ling Huang
arXiv Machine Learning
Sep 16

Evaluating Open-Weight E-Commerce Agents with Environment-Grounded Verification

The paper introduces a deterministic, reproducible e‑commerce environment that pre‑commits customer and trajectory parameters, enabling a simulated consumer to attempt purchasing a target cart with the help of an evaluated model. The environment records every assistant action and state, allowing post‑trial evaluation of specific conversation components and applying penalties based on tool‑call accuracy. Using this setup, the authors benchmark eight open‑weight agents (20B–35B parameters) across 160 trials and 44 metrics, revealing nuanced performance issues such as under‑action, over‑purchase, unsupported product attributes, and poor search that are hidden by overall success rates.

By Nimit Shah, Haitz S\'aez de Oc\'ariz Borde
arXiv AI
Aug 20

Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-Engagement

The paper introduces a production-ready framework that connects e‑commerce search and CRM systems via AI‑powered Product Research Agents. These agents detect users with exploratory purchase intent, perform multi‑agent research using behavioral data, external knowledge, and catalog information, and then send personalized product recommendations through WhatsApp. In a 23‑day deployment, the system sent about 15,000 notifications, achieving higher click‑through rates than standard campaigns and generating downstream purchases and GMV gains.

By Mandar Kulkarni, Pooja A., Samir Shah