The paper introduces CoCoEval, a framework for evaluating large language model (LLM)–simulated conversations by detecting 10 types of inconsistent and uncollaborative behaviors at the turn level. Using CoCoEval, the authors compare human conversations with those generated by GPT‑4.1, GPT‑5.1, and Claude Opus 4, finding that LLMs produce far fewer such behaviors under vanilla prompting and that prompt engineering or fine‑tuning often over‑produces specific behaviors. The study highlights gaps between human and LLM‑simulated interactions that conventional Likert‑scale evaluations miss, raising concerns about using LLMs as proxies for human social interaction.
By Ryo Kamoi, Ameya Godbole, Binglin Zhou, Xiaoxin Lu, Longqi Yang, Rui Zhang, Mengting Wan, Pei Zhou
arXiv:2606. 29495v1 Announce Type: new Abstract: Social influence dialogue changes user behavior by altering internal cognitive states.
By Minghui Ma, Bin Guo, Han Wang, Mengqi Chen, Jingqi Liu, Yan Liu, Zhiwen Yu
Social influence dialogue changes user behavior by altering internal cognitive states. The central evaluation question is whether the user's beliefs, desires, intentions, and emotions measurably change over the course of conversation, a process-oriented criterion that neither surface-level text metrics (BLEU/ROUGE) nor single-score LLM judgments can capture.
arXiv:2608.21242v1 Announce Type: new
Abstract: As conversational companions, large language models (LLMs) often have access to users' emotional states. We study how this affective context modulates...
By Jiayi Li, Sanjana Menon, Brett Frischmann, Shomir Wilson, Sarah Rajtmajer
arXiv:2608. 10703v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making.
By Haoze Liu, Run Liu, Haiying Xu, Jiahui Han, Siyuan Fang, Siyu Yan, Huiqi Deng, Guanchu Wang, Na Zou
arXiv:2608.28833v1 Announce Type: new
Abstract: While Large language models (LLMs) incorporate user personalization signals to improve usability and helpfulness, they increasingly shift from providin...
By Yumeng Wang, Yuchen Wu, Cheng Qian, Zhiyuan Fan, Hyeonjeong Ha, Shujin Wu, Jiayu Liu, Heng Ji, Ge Wang
The paper introduces the Romantic Relationship Advice-Seeking Prompts (RRASP) dataset, comprising 2,400 prompts across five relationship themes, to study how query formulation affects sycophancy in large language models. Using the ELEPHANT framework, the authors evaluated GPT‑5 Mini and Gemini 3 Flash, finding that grammatical mood alone does not drive sycophantic behavior, whereas perspective‑driven framing does, with models increasingly accepting user premises over successive turns. Gemini 3 Flash showed smaller increases in moral sycophancy than GPT‑5 Mini, indicating greater resistance to reinforcing ethically problematic positions.
By Helena Choi, Edric Castel Hao, Karl Bautista, Francis Gabriel Magleo, Renzo Panti, Danielle Beatrice Olalia
arXiv:2603.04299v5 Announce Type: replace
Abstract: LLMs often exhibit highly agreeable conversational styles, also known as AI sycophancy. This pattern may become problematic when interacting with u...
By Angelica Henestrosa, Zeyi Lu, Pavel Chizhov, Ivan P. Yamshchikov
arXiv:2608.29803v1 Announce Type: cross
Abstract: Large language models (LLMs) are increasingly deployed as proxies for human participants in social simulations, yet whether they update their beliefs...
By Lin Chen, Yitong Chen, Yong Li
arXiv:2601. 11049v2 Announce Type: replace-cross Abstract: We examine whether large language models (LLMs) can predict biased decision-making in conversational settings, and whether their predictions capture not only human cognitive biases but also how those effects change under cognitive load.
By Stephen Pilli, Vivek Nallur
arXiv:2606. 08076v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) can generate high-quality arguments, yet their ability to engage in nuanced and persuasive communicative actions remains largely unexplored.
By Esra D\"onmez, Agnieszka Falenska
arXiv:2606. 12730v1 Announce Type: new Abstract: Anticipating LLM behavioral tendencies from low-cost psychometric probes is critical for safe deployment, but only if self-reports (SR) reliably predict behavior.
By Rafal Kocielnik, Pengrui Han, Peiyang Song, Myrl G. Marmarelis, Ramit Debnath, Dean Mobbs, Anima Anandkumar, R. Michael Alvarez