arXiv:2606. 14715v1 Announce Type: cross Abstract: LLM agents are increasingly used to simulate real world interactions, but it remains unclear whether simulated behaviors preserve the content patterns and interaction dynamics of real human behaviors.
By Yaoning Yu, Ye Yu, Haojing Luo, Haohan Wang
arXiv:2606.06443v3 Announce Type: replace
Abstract: Large language models are increasingly used to simulate social media users and infer how individuals may respond to online discussions. However, it...
By Xinnong Zhang, Wanting Shan, Hanjia Lyu, Zhongyu Wei, Jiebo Luo
arXiv:2601.12208v2 Announce Type: replace
Abstract: Evaluating conversational systems in multi-turn settings remains a fundamental challenge. Conventional pipelines typically rely on manually defined...
By Yunzhe Li, Richie Yueqi Feng, Tianxin Wei, Chin-Chia Hsu
The paper introduces a hypothesis-driven simulation workflow that screens customer experience (CX) agents before deployment, using synthetic customers and simulated tool outputs to emulate multi-step interactions without accessing production backends. Applied to Nubank’s high-volume Card Delivery and Card Management chat-support agents, the simulation’s binary evaluator scores correlated strongly with production results, and simulation-guided iterations raised transactional net promoter score by 36.69 points in a live A/B test. Additionally, screening over 16,000 simulated conversations helped select a model that increased self‑service rate by 8.82 percentage points without harming net promoter score, demonstrating that simulation enables extensive model exploration safely.
By Edesio Alcoba, Kevin Rossell, Aman Gupta, Shao Tang, Jiwoo Hong, Pabel Carrillo-Mendoza, Wanderson Concei\c{c}\~ao Ferreira, Alvaro Tedeschi, Zayd Simjee, Shreya Rajpal, Bruno Finardi Hime, Christian Sousa, Luis Moneda, Herbert Fei, Daniel Silva, Rohan Ramanath
arXiv:2606. 06027v1 Announce Type: cross Abstract: Community-conditioned language model adaptation requires choices about data collection, community definition, and evaluation that are currently made independently in each study, making it hard to compare assumptions or reuse artifacts.
By Amirhossein Ghaffari, Ali Goodarzi, Huong Nguyen, Simo Hosio, Lauri Lov\'en, Ekaterina Gilman
arXiv:2607. 25218v1 Announce Type: new Abstract: Debt collection is a critical negotiation task in the financial industry, with strong practical relevance and exceptional academic value as a behaviorally rich, high-stakes testbed for human-centered dialogue systems.
By Yuhang Yang, Kai Tang, Chao Ye, Haobo Wang, Qiqi Luo, Jinguang Zheng, Zhixin Zhang
arXiv:2510.05124v3 Announce Type: replace
Abstract: We propose MADS (Multi-Agent Dialogue Simulation), a scalable framework for generating persuasive multi-turn dialogues via agent self-play. MADS em...
By Mingjin Li, Yu Liu, Huayi Liu, Xiang Ye, Chao Jiang, Hongguang Zhang, Yu Ruan
DocuTeam is a mixed‑initiative multi‑agent discussion system that allows both users and agents to start and steer conversations around evolving documents. Agents monitor changes to the document and proactively initiate or redirect discussions, while users can shape the dialogue or adopt agent suggestions. In a within‑subjects study with 20 participants, DocuTeam produced outcomes that were rated as more novel, relevant, and specific compared to a baseline, without increasing cognitive load.
By Heechan Lee, Juhyeon Choi, Tae Soo Kim, Juho Kim, Joseph Seering
TxSum introduces a user-centered approach to understanding Ethereum transactions by providing structured, risk-aware explanations grounded at the token‑flow level. The authors built a dataset of 187 complex transactions with 2,375 token‑flow annotations and transaction‑level summaries, and developed MATEX, a multi‑agent framework that retrieves external knowledge and audits explanations for factual consistency. MATEX outperforms existing baselines, improving user comprehension from 52.9% to 76.5% and increasing malicious‑transaction rejection from 36.0% to 88.0% while keeping false‑rejection rates low.
By Zifan Peng, Jingyi Zheng, Yule Liu, Huaiyu Jia, Qiming Ye, Jingyu Liu, Xufeng Yang, Mingchen Li, Qingyuan Gong, Xuechao Wang, Xinlei He
arXiv:2606. 05178v1 Announce Type: cross Abstract: As AI-driven product development accelerates, the bottleneck is shifting from how we build to what we build.
By Tim Dorn, Saara A. Khan, Julie Mumford
arXiv:2606. 18268v1 Announce Type: cross Abstract: Community-based fact-checking that relies on cross-consensus is expanding rapidly on social media platforms.
By Changxi Wen, Shuning Zhang, Bohao Chu, Yuwei Chuai, Hui Wang, Dai Shi, Xin Yi, Hewu Li
arXiv:2606. 05890v1 Announce Type: cross Abstract: LLMs are increasingly deployed as Artificial Moral Advisors (AMA) in a variety of contexts: what kind of conversational patterns should they display?
By Salvatore Greco, Hainiu Xu, Jacopo Domenicucci, Yulan He, Sylvie Delacroix