Beyond Cooperative Simulators: Generating Realistic User Personas for Robust Evaluation of LLM Agents
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The paper introduces Calibrated User Embeddings (CUE), a framework that encodes real user sessions into continuous representations and decodes them into persona commands to steer large language models (LLMs) as user simulators without additional training. CUE enables user-conditioned replay of past interactions and generates novel personas that better align with real-user success rates and failure patterns. Evaluations on the τ²‑Bench dataset show that CUE‑driven simulators commit fewer simulator‑attributed errors, more accurately reproduce real‑user failure modes, and maintain competitive user fidelity across various tasks and LLMs.
PersonaForge is a user‑simulation framework that generates realistic multi‑turn interactions between users and agentic systems, addressing the gap that most training data assumes single‑turn queries. It uses a four‑dimensional persona space, SOUL‑driven behavioral control calibrated to real‑user statistics, and Reverse Deep Construction from authentic seed queries to create a 6.3K‑record training set and a 138‑task benchmark called PersonaForge‑Bench across 20 professional domains. Experiments with Qwen3.5‑27B show that training with PersonaForge improves composite scores by 4.1%, especially in Task Completion (+6.0%) and Response Quality (+6.8%), while also reducing turns and tool calls, indicating more efficient interactions.
arXiv:2507. 09788v3 Announce Type: replace-cross Abstract: Recent advances in Large Language Models (LLM) have led to a new class of autonomous agents, renewing and expanding interest in the area.
The paper introduces a three-tier persona vector for user simulation in evaluating LLM agents, comprising 23 dimensions across demographics, behavioral traits, and emotional states, plus a query-complexity overlay. It demonstrates that these nuanced personas generate diverse, scenario-reactive conversations, leading to significant variations in agent goal achievement and compliance across different contexts. The model’s design allows for reproducible, auditable user behavior patterns without relying on learned covariance matrices.
arXiv:2608. 19549v1 Announce Type: new Abstract: This paper addresses the issue of the significant labor required to test interview dialogue systems.
arXiv:2510. 04491v3 Announce Type: replace Abstract: Despite rapid progress in building conversational AI agents, robustness is still largely untested.