arXiv:2606. 18258v1 Announce Type: cross Abstract: Large language models (LLMs) exhibit a wide range of human-like behaviors, from expressing thoughts and emotions, to engaging in relationship-building with users, to refusing requests and maintaining boundaries.
By Sunnie S. Y. Kim, Margit Bowler, Leon A Gatys
arXiv:2607. 20734v1 Announce Type: new Abstract: As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction.
By Jihoon Tack, Philippe Laban, Jennifer Neville
arXiv:2601.12208v2 Announce Type: replace
Abstract: Evaluating conversational systems in multi-turn settings remains a fundamental challenge. Conventional pipelines typically rely on manually defined...
By Yunzhe Li, Richie Yueqi Feng, Tianxin Wei, Chin-Chia Hsu
arXiv:2607. 12085v1 Announce Type: new Abstract: Evaluating retail conversational agents requires methods beyond lexical-overlap metrics to assess intent alignment, factuality, helpfulness, clarity, tone, and overall response quality.
By Niranjan Kumar M, Balaji Nagarajan, Karthik Nair, Faysal Satter, Nithin Surendran
PRAGMA is a benchmark designed to evaluate personalized guidance in long‑term conversations. It includes curated longitudinal conversation histories, evidence annotations, and guidance scenarios that reflect evolving user contexts and incorrect assumptions. Experiments show that current retrieval, memory, and long‑context models struggle to recover relevant conversational evidence and to use it effectively for personalized guidance.
By Hyojeong Yu, Hyukhun Koh, Minsung Kim, Yunah Jang, Kyomin Jung
The paper introduces CoCoEval, a framework for evaluating large language model (LLM)–simulated conversations by detecting 10 types of inconsistent and uncollaborative behaviors at the turn level. Using CoCoEval, the authors compare human conversations with those generated by GPT‑4.1, GPT‑5.1, and Claude Opus 4, finding that LLMs produce far fewer such behaviors under vanilla prompting and that prompt engineering or fine‑tuning often over‑produces specific behaviors. The study highlights gaps between human and LLM‑simulated interactions that conventional Likert‑scale evaluations miss, raising concerns about using LLMs as proxies for human social interaction.
By Ryo Kamoi, Ameya Godbole, Binglin Zhou, Xiaoxin Lu, Longqi Yang, Rui Zhang, Mengting Wan, Pei Zhou