arXiv:2608. 09939v1 Announce Type: cross Abstract: Production teams deploying LLM chat agents face a specific quality assurance gap: existing evaluation tools test individual responses or simulate social interactions, but none systematically verify whether real users can achieve their goals through multi-turn conversation.
By Alexandre Cristov\~ao Maiorano
The paper introduces a deterministic, reproducible e‑commerce environment that pre‑commits customer and trajectory parameters, enabling a simulated consumer to attempt purchasing a target cart with the help of an evaluated model. The environment records every assistant action and state, allowing post‑trial evaluation of specific conversation components and applying penalties based on tool‑call accuracy. Using this setup, the authors benchmark eight open‑weight agents (20B–35B parameters) across 160 trials and 44 metrics, revealing nuanced performance issues such as under‑action, over‑purchase, unsupported product attributes, and poor search that are hidden by overall success rates.
By Nimit Shah, Haitz S\'aez de Oc\'ariz Borde
arXiv:2603. 29888v2 Announce Type: replace-cross Abstract: In collaboration with Alibaba, we study how a generative AI assistant affects service performance in e-commerce after-sales operations.
By Xiao Ni, Yiwei Wang, Tianjun Feng, Lauren Xiaoyan Lu, Yitong Wang, Congyi Zhou
The paper presents an AI‑powered triaging agent for banks that uses large language models to conduct multi‑turn conversations, ask relevant questions, and classify customer cases for accurate routing to specialist teams. The system is integrated with policy, safety guardrails, and reasoning frameworks, and its performance is evaluated using synthetic digital twins that simulate realistic, labeled dialogues based on historical data. Results show a 30.6% increase in classification accuracy and high satisfaction from subject‑matter experts, demonstrating the effectiveness of targeted probing for scalable banking operations.
By Alankar Atreya, Stefan Sylvius Wagner, Devesh Batra, Robert Hankache, Cristovao Iglesias Jr, Patrick Sinclair, Giulio Pelosio, Michael McMillan, Greig A. Cowan, Raad Khraishi
arXiv:2606. 12924v1 Announce Type: new Abstract: We present a modular two-agent simulation framework for evaluating conversational shopping assistant architectures.
By Jetlir Duraj, Jayanth Yetukuri, Shuang Zhou, Dhruv Varma, Rui Kong, Ishita Khan, Qunzhi Zhou
arXiv:2608. 04205v1 Announce Type: new Abstract: Human evaluation of AI systems and digital products is costly, slow, and difficult to scale.
By Xiaomin Li, Yuexing Hao, Jianheng Hou, Jintao Huang, Qianfeng Wen, Shirley Huang, Yifan Liu, Xiaoyi Liu, Yilan Fan, Yijun Wang, Koutian Wu, Ruoqi Gao, Muhammad Ahmed Mohsin, Jing Tang, Brihi Joshi, Heming Liu, Zheyuan Deng, Zonglin Di, Sankalp Jajee, Jiuyao Lu, Zhiwei Zhang, Saksham Kapoor, Ishan Gupta, Yunhan Zhao, Chanwoo Park, Yucheng Lu, Bing Hu, Weihang Xiao, Aravind Mohan, Hanwen Xing, Runyu Zhang, Mihir Kulshreshtha, Yuanda Xu, Qianyu Zhu, Dianzhuo Wang, Yuxin Xiao, Bowen Jiang, Yongye Su, Wenhao Chai, Zuxin Liu, Lawrence Yunliang Chen, Xuandong Zhao, Ethan Ye, Shivam Patel, Jason Xie, Alex Martin Richmond, Weixiang Ding, Emre Okcular, Diya Mathew, Ziheng Wang, Rana M. Shahroz Khan, Zhejian Peng, Fang Wu, Fan Nie, Xinyang Han, Yubin Kim, Jiawei Zhang, Zhenting Qi, Huangyuan Su, Xu Pan, Abinitha Gourabathina, Hyewon Jeong, Hemanth Neelgund Ramesh, Kumail Alhamoud, Kimia Hamidieh, Zidi Xiong, Samuel Schmidgall, Pengrui Han, Yepeng Huang, Yongheng Wang, Bowen Yang, Alex Gu, Yuchu Wang, Akshay Paruchuri, Brenna Li, Hejie Cui, Jiayuan Ding, Chaosheng Dong, Jiahao Wang, Yixuan He, Chi Wang, Pamela Bhattacharya, Tianyi Peng, Paul Pu Liang, Mitchell Gordon, Yilun Du, Marinka Zitnik, James Zou, Prasanna Tambe, Philip Torr, Emily Fox, Asu Ozdaglar, Dawn Song
The paper introduces a three-tier persona vector for user simulation in evaluating LLM agents, comprising 23 dimensions across demographics, behavioral traits, and emotional states, plus a query-complexity overlay. It demonstrates that these nuanced personas generate diverse, scenario-reactive conversations, leading to significant variations in agent goal achievement and compliance across different contexts. The model’s design allows for reproducible, auditable user behavior patterns without relying on learned covariance matrices.
By Rahul Khedar, Eshita, Sneha Teja Sree Reddy Thondapu, Mayank Malhotra, Arup Kumar Das, Jitesh Chandra Mishra, Arun Menon, Avinash Karn, Mouli V
The paper introduces a three-tier persona vector to generate diverse, realistic user inputs for evaluating tool-augmented LLM agents. The vector includes 23 dimensions: categorical demographics, continuous behavioral traits, and continuous emotional states, plus a query-complexity overlay. Experiments on 64,698 conversations show that these persona dimensions produce measurable differences in agent performance and realistic scenario-reactive behavior.
arXiv:2607. 28956v1 Announce Type: new Abstract: Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria.
By Qiming Shi, Yulong Tao, Linbo Jin, Zhaolu Kang, Yibo Dou, Jiawen Zhu, Tianjun Pan, Shaokang Fu, Chengyu Wang, Siyue Li, Yaping Cheng, Di Weng, Chengfu Huo
arXiv:2512. 04123v4 Announce Type: replace-cross Abstract: LLM-based agents already operate in production across many industries, yet we lack an understanding of what technical methods make deployments successful.
By Melissa Z. Pan, Negar Arabzadeh, Riccardo Cogo, Yuxuan Zhu, Alexander Xiong, Lakshya A Agrawal, Huanzhi Mao, Emma Shen, Sid Pallerla, Liana Patel, Shu Liu, Tianneng Shi, Xiaoyuan Liu, Jared Quincy Davis, Emmanuele Lacavalla, Alessandro Basile, Shuyi Yang, Paul Castro, Daniel Kang, Koushik Sen, Dawn Song, Joseph E. Gonzalez, Ion Stoica, Matei Zaharia, Marquita Ellis
The paper describes a production migration of a large-scale customer‑support conversational assistant from a single blended model to a Dynamic Response (DR) architecture. DR replaces the Qwen3‑235B‑A22B responder with a bounded ReAct orchestrator that selects typed tools and a smaller generator that writes from a validated context contract. The migration yields significant improvements: precision‑first entity selection boosts reservation selector precision from 8.3% to 89.1%, typed action IDs eliminate structured‑action hallucination, and hard‑escalation responses drop from 5.60% to 3.08%. Latency is reduced from 3.87 s to 2.24 s, GPU usage is cut by roughly one‑third, and self‑hosting cuts annual model‑serving costs by more than an order of magnitude.
By Cen Mia Zhao, Peng Wang, Chuan Shi, Yufeng Zhang, Ying Lyu, Wanmeng Ren, Robert Xue, Claire Na Cheng, Yashar Mehdad
The paper introduces AgentX-Model, a dual‑agent framework that links proposal development with model experimentation in industrial recommender systems. The Research Agent drafts proposals from literature and prior findings, while the Model Agent runs multi‑round experiments, returning code, metrics, and open questions. The framework iteratively selects starting implementations and formulates new research questions, organizing work into Reproduce, Follow‑up, Composition, and Diagnose actions. Across production evaluations, most experiments exceeded business baselines, with recent A/B tests showing significant gains in acquisition efficiency, advertising spend, and watch time while reducing computational cost.
By Shuang Yang, Zijie Zhuang, Changxin Lao, Pengbo Xu, Hanwen Xu, Yusheng Huang, Han Gao, Guanchen Wang, Tianbao Ma, Linxun Chen, Peilin Song, Xuming Wang, Chen Li, Fan Wu, Tao Wang, Zibo Zhao, Xiangyu Wu, An Liu, Fei Pan, Peng Jiang, Chen Yang, Zhaojie Liu, Wenwu Ou