arXiv AI

Generative AI and Sales Productivity: Field Experiments in Online Retail

arXiv:2510. 12049v4 Announce Type: replace-cross Abstract: We quantify the short-term impact of Generative Artificial Intelligence (GenAI) on sales performance through a series of large-scale randomized field experiments involving millions of users and products at a leading cross-border online retail platform.

arXiv Machine Learning
Sep 1

E-Commerce Bench: Evaluating LLM Agents on Long-Horizon Autonomous Business Operation

E-Commerce Bench is an open‑source benchmark that simulates a year‑long e‑commerce operation, requiring LLM agents to manage multiple online stores, negotiate with suppliers, optimize sales, fulfill orders, handle returns, and manage cash flow. The environment uses real product and supplier data, a calendar of promotions and shocks, and deterministic customer and negotiation models to enable reproducible evaluation. The study evaluates 18 state‑of‑the‑art models across seven metrics, finding no single model dominates, with GPT‑5.6 Sol achieving the highest year‑end assets but lagging in fraud avoidance and operational efficiency.

By Wei Fan, Xinjie Shen, Xudong Guo, Jianhong Tu, Yang Su, Yinger Zhang, Lianghao Deng, Fengyu Wang, Baohua Dong, Yangqiu Song, Dayiheng Liu
arXiv AI
Aug 20

Bridging Search and CRM: Productionizing AI Product Research Agents for Customer Re-Engagement

The paper introduces a production-ready framework that connects e‑commerce search and CRM systems via AI‑powered Product Research Agents. These agents detect users with exploratory purchase intent, perform multi‑agent research using behavioral data, external knowledge, and catalog information, and then send personalized product recommendations through WhatsApp. In a 23‑day deployment, the system sent about 15,000 notifications, achieving higher click‑through rates than standard campaigns and generating downstream purchases and GMV gains.

By Mandar Kulkarni, Pooja A., Samir Shah
arXiv AI
Sep 24

Shopping by algorithm: How agentic AI deploys human heuristics as a surrogate consumer

The study investigates how Large Language Models (LLMs) acting as surrogate consumers are influenced by marketing pricing cues such as just‑below pricing and promotional framing. Using a tool called "Tool‑Lab" to trace information acquisition, the researchers found that when no cost is imposed, pricing cues rarely mislead LLMs, but when acquisition costs are introduced under a vague goal prompt, LLMs tend to omit important diagnostic attributes and make suboptimal choices similar to human heuristics. The findings suggest that marketing heuristics in AI‑driven shopping are shaped more by storefront information architecture than by inherent LLM limitations.

By Davood Wadi, Yu Ma
arXiv Machine Learning
Aug 27

DCEO: Direct Causal Effect Optimization for Long-Term User Value Modeling in E-commerce Search

The paper introduces DCEO, a data‑driven framework that learns item‑level proxy scores directly aligned with long‑term user objectives in e‑commerce search. It aggregates these scores into a user‑level metric, measures alignment via relative causal effect, and uses an actor‑critic model to generate context‑dependent fusion weights for multiple objectives. Offline experiments and a 41‑day online A/B test show DCEO improves GMV by 0.36% over traditional proxies.

By Junzhao Zhang, Tao Zhang, Liren Yu, Feiyi Dong, Zhixuan Zhang, Dan Ou, Haihong Tang