The paper introduces Calibrated User Embeddings (CUE), a framework that encodes real user sessions into continuous representations and decodes them into persona commands to steer large language models (LLMs) as user simulators without additional training. CUE enables user-conditioned replay of past interactions and generates novel personas that better align with real-user success rates and failure patterns. Evaluations on the τ²‑Bench dataset show that CUE‑driven simulators commit fewer simulator‑attributed errors, more accurately reproduce real‑user failure modes, and maintain competitive user fidelity across various tasks and LLMs.
By Anjali Kantharuban, Jonas Mueller
PersonaForge is a user‑simulation framework that generates realistic multi‑turn interactions between users and agentic systems, addressing the gap that most training data assumes single‑turn queries. It uses a four‑dimensional persona space, SOUL‑driven behavioral control calibrated to real‑user statistics, and Reverse Deep Construction from authentic seed queries to create a 6.3K‑record training set and a 138‑task benchmark called PersonaForge‑Bench across 20 professional domains. Experiments with Qwen3.5‑27B show that training with PersonaForge improves composite scores by 4.1%, especially in Task Completion (+6.0%) and Response Quality (+6.8%), while also reducing turns and tool calls, indicating more efficient interactions.
By Hanglong Lv, Dawei Zhu, Lei Li, Bowen Ye, Huaqiu Liu, Yifan Song, Bofei Gao, Weimin Xiong, Jinhao Dong, Chenhong He, Lingpeng Kong, Qi Liu, Tong Yang, Fuli Luo
arXiv:2507. 09788v3 Announce Type: replace-cross Abstract: Recent advances in Large Language Models (LLM) have led to a new class of autonomous agents, renewing and expanding interest in the area.
By Paulo Salem, Robert Sim, Christopher Olsen, Prerit Saxena, Rafael Barcelos, Yi Ding
The paper introduces a three-tier persona vector for user simulation in evaluating LLM agents, comprising 23 dimensions across demographics, behavioral traits, and emotional states, plus a query-complexity overlay. It demonstrates that these nuanced personas generate diverse, scenario-reactive conversations, leading to significant variations in agent goal achievement and compliance across different contexts. The model’s design allows for reproducible, auditable user behavior patterns without relying on learned covariance matrices.
By Rahul Khedar, Eshita, Sneha Teja Sree Reddy Thondapu, Mayank Malhotra, Arup Kumar Das, Jitesh Chandra Mishra, Arun Menon, Avinash Karn, Mouli V
arXiv:2608. 19549v1 Announce Type: new Abstract: This paper addresses the issue of the significant labor required to test interview dialogue systems.
By Mikio Nakano, Kazunori Komatani, Hironori Takeuchi
arXiv:2510. 04491v3 Announce Type: replace Abstract: Despite rapid progress in building conversational AI agents, robustness is still largely untested.
By Muyu He, Anand Kumar, Tsach Mackey, Meghana Rajeev, James Zou, Nazneen Rajani
The paper introduces a three-tier persona vector to generate diverse, realistic user inputs for evaluating tool-augmented LLM agents. The vector includes 23 dimensions: categorical demographics, continuous behavioral traits, and continuous emotional states, plus a query-complexity overlay. Experiments on 64,698 conversations show that these persona dimensions produce measurable differences in agent performance and realistic scenario-reactive behavior.
arXiv:2606. 14199v1 Announce Type: cross Abstract: Large language models are increasingly deployed as human simulators for interactive evaluation and social simulation.
By Xuhui Zhou, Weiwei Sun, Weihua Du, Jiarui Liu, Haojia Sun, Qianou Ma, Tongshuang Wu, Yiming Yang, Maarten Sap
arXiv:2606. 18263v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to simulate human populations via persona prompting, often under the assumptions that richer persona descriptions improve behavioral fidelity, similarly sized attribute combinations are equally simulatable, and persona definitions generalize across tasks.
By Aanisha Bhattacharyya, Yaman Kumar Singla, Rajiv Ratn Shah, Changyou Chen, Jitendra Ajmera
arXiv:2609.22607v1 Announce Type: new
Abstract: We argue here that the current dominant practice in LLM human simulation: prompting instruction-tuned assistant language models to role-play personas,...
By Minwoo Kang, T\'ea Wright, Seun Eisape, Ayush Raj, Suhong Moon, Joseph Suh, Alane Suhr, David M. Chan, John Canny
PERSONAWEAVER is a new approach to procedural character generation that separates world building from behavioral specification, using manually curated banks of moral positions and conversational reactions to diversify character behavior. By applying this method across ten realistic and fantastical settings and three large language models, the system produces broader moral and interactional response distributions, varied interpersonal language, response length, sentiment, and less archetypal world attribute combinations compared to prior work.
By Maan Qraitem, Kate Saenko, Bryan A. Plummer
arXiv:2602. 12394v2 Announce Type: replace Abstract: Personalized prompting offers large opportunities for deploying large language models (LLMs) to diverse users, yet existing prompt optimization methods primarily focus on task-level optimization while largely overlooking user-specific preferences and latent constraints of individual users.
By Yuchen Ma, Yue Huang, Wenjie Wang, Xiaonan Luo, Xiangliang Zhang, Stefan Feuerriegel