arXiv:2603. 16453v3 Announce Type: replace Abstract: Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments remains uncertain.
By Linghua Zhang, Jun Wang, Jingtong Wu, Zhisong Zhang
arXiv:2606. 15862v1 Announce Type: new Abstract: Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments remains uncertain.
By Linghua Zhang, Jun Wang, Jingtong Wu, Zhisong Zhang
The paper introduces a deterministic, reproducible e‑commerce environment that pre‑commits customer and trajectory parameters, enabling a simulated consumer to attempt purchasing a target cart with the help of an evaluated model. The environment records every assistant action and state, allowing post‑trial evaluation of specific conversation components and applying penalties based on tool‑call accuracy. Using this setup, the authors benchmark eight open‑weight agents (20B–35B parameters) across 160 trials and 44 metrics, revealing nuanced performance issues such as under‑action, over‑purchase, unsupported product attributes, and poor search that are hidden by overall success rates.
By Nimit Shah, Haitz S\'aez de Oc\'ariz Borde
arXiv:2605. 07699v2 Announce Type: replace-cross Abstract: LLM-based agents are increasingly deployed for routine but consequential tasks in real-world domains, where their behavior is governed by inherently ambiguous domain policies that admit multiple valid interpretations.
By Hsuvas Borkakoty, Sebastian Pohl, Cheng Wang, Bei Chen, Yufang Hou
arXiv:2606. 16183v1 Announce Type: cross Abstract: We develop an LLM-powered virtual population model that simulates demand for pricing decisions, in settings where products are described by rich unstructured information, such as text descriptions and images, and where decision makers need not only mean-demand predictions but also uncertainty estimates for counterfactual prices.
By Chengpiao Huang, Kaizheng Wang
arXiv:2607. 28956v1 Announce Type: new Abstract: Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria.
By Qiming Shi, Yulong Tao, Linbo Jin, Zhaolu Kang, Yibo Dou, Jiawen Zhu, Tianjun Pan, Shaokang Fu, Chengyu Wang, Siyue Li, Yaping Cheng, Di Weng, Chengfu Huo
The paper introduces SalesLLM, a bilingual (Chinese/English) benchmark for evaluating large language models (LLMs) in realistic sales dialogues. It comprises 30,074 scripted configurations and 1,805 curated multi‑turn scenarios from Financial Services and Consumer Goods, with controllable difficulty and personas. An automatic evaluation pipeline uses an LLM judge for sales‑process progress and fine‑tuned BERT classifiers for end‑of‑dialogue buying intent, while a user model, CustomerLM, is trained to improve simulation fidelity. SalesLLM scores correlate strongly with human ratings (Pearson r = 0.86) and reveal that top Chinese LLMs match junior‑to‑intermediate human salespeople but not experts, with cross‑lingual consistency remaining poor.
By Xuanbo Su, Wenhao Hu, Le Zhan, Yuting Xie, Kailin Lyu, Kaijie Chen, Ziwei Li, Yeqiang Wang, Haibo Su, Yunzhang Chen, Ling Huang
The paper introduces a hypothesis-driven simulation workflow that screens customer experience (CX) agents before deployment, using synthetic customers and simulated tool outputs to emulate multi-step interactions without accessing production backends. Applied to Nubank’s high-volume Card Delivery and Card Management chat-support agents, the simulation’s binary evaluator scores correlated strongly with production results, and simulation-guided iterations raised transactional net promoter score by 36.69 points in a live A/B test. Additionally, screening over 16,000 simulated conversations helped select a model that increased self‑service rate by 8.82 percentage points without harming net promoter score, demonstrating that simulation enables extensive model exploration safely.
By Edesio Alcoba, Kevin Rossell, Aman Gupta, Shao Tang, Jiwoo Hong, Pabel Carrillo-Mendoza, Wanderson Concei\c{c}\~ao Ferreira, Alvaro Tedeschi, Zayd Simjee, Shreya Rajpal, Bruno Finardi Hime, Christian Sousa, Luis Moneda, Herbert Fei, Daniel Silva, Rohan Ramanath
arXiv:2607. 06993v1 Announce Type: new Abstract: Customer behavior modeling underpins recommendation, marketing, and decision support, yet existing approaches either optimize predictive accuracy without explaining decisions or simulate users without grounding them in real behavioral data.
By Wachiravit Modecrua, Krittin Pachtrachai, Touchapon Kraisingkorn
arXiv:2510. 12049v4 Announce Type: replace-cross Abstract: We quantify the short-term impact of Generative Artificial Intelligence (GenAI) on sales performance through a series of large-scale randomized field experiments involving millions of users and products at a leading cross-border online retail platform.
By Lu Fang, Zhe Yuan, Kaifu Zhang, Dante Donati, Miklos Sarvary
The paper introduces a simulation framework that uses large language model (LLM) agents conditioned on data‑driven personas to predict A/B test outcomes. These personas are built from anonymized user behavioral patterns, engagement signals, and inferred demographics, offering a more realistic population model than synthetic or rule‑based personas. The authors evaluate question design, persona data source, behavioral depth versus diversity, and population subsampling, achieving 0.75–0.90 directional accuracy on 40 real A/B tests.
By Ziyad Benomar, Weronika {\L}ajewska, Leonardo Perelli, Saab Mansour
E-Commerce Bench is an open‑source benchmark that simulates a year‑long e‑commerce operation, requiring LLM agents to manage multiple online stores, negotiate with suppliers, optimize sales, fulfill orders, handle returns, and manage cash flow. The environment uses real product and supplier data, a calendar of promotions and shocks, and deterministic customer and negotiation models to enable reproducible evaluation. The study evaluates 18 state‑of‑the‑art models across seven metrics, finding no single model dominates, with GPT‑5.6 Sol achieving the highest year‑end assets but lagging in fraud avoidance and operational efficiency.
By Wei Fan, Xinjie Shen, Xudong Guo, Jianhong Tu, Yang Su, Yinger Zhang, Lianghao Deng, Fengyu Wang, Baohua Dong, Yangqiu Song, Dayiheng Liu