The paper investigates whether large language model (LLM) agents can simulate public deliberation by reflecting population opinion patterns and producing interaction-driven opinion change. Using census‑grounded Korean personas debating real policy questions, the study finds that persona agents fail to reliably reproduce population opinion patterns, often concentrating responses and reversing demographic differences. While deliberations generate reasoned, reciprocal arguments and some stance movement, much of this change occurs without peer exchange, and anchoring agents to population‑informed starting positions suppresses updating, indicating that population representation, argument generation, and interaction‑driven opinion change do not necessarily align.
By Chaemin Jang, Junsik Min, Jaewoo Choi, Donggyu Lee, Haiin Lee, Junyoung Park, Namhee Kim, Hyunwoo Kim, Jungwon Kim, Juho Kim, Nuri Kim, Jihee Kim
arXiv:2609.22607v1 Announce Type: new
Abstract: We argue here that the current dominant practice in LLM human simulation: prompting instruction-tuned assistant language models to role-play personas,...
By Minwoo Kang, T\'ea Wright, Seun Eisape, Ayush Raj, Suhong Moon, Joseph Suh, Alane Suhr, David M. Chan, John Canny
The paper introduces CAPA, a Collaborative Agent Predictive Architecture designed to give large language model (LLM) agents situational awareness in online meetings. CAPA uses a Perceiver to update meeting state, a Predictor to forecast conversation flow, a Controller to decide speaking actions, and a Generator to phrase contributions. Evaluated on 137 AMI meetings, CAPA reduces the silence rate from 51.4% to 2.5%, doubles credited recovery, and maintains low hallucination, demonstrating that structured state tracking is key to effective delegation.
By Muneeb Khan, Frederic Kirstein, Terry Ruas, Bela Gipp
The paper introduces CoCoEval, a framework for evaluating large language model (LLM)–simulated conversations by detecting 10 types of inconsistent and uncollaborative behaviors at the turn level. Using CoCoEval, the authors compare human conversations with those generated by GPT‑4.1, GPT‑5.1, and Claude Opus 4, finding that LLMs produce far fewer such behaviors under vanilla prompting and that prompt engineering or fine‑tuning often over‑produces specific behaviors. The study highlights gaps between human and LLM‑simulated interactions that conventional Likert‑scale evaluations miss, raising concerns about using LLMs as proxies for human social interaction.
By Ryo Kamoi, Ameya Godbole, Binglin Zhou, Xiaoxin Lu, Longqi Yang, Rui Zhang, Mengting Wan, Pei Zhou
arXiv:2606. 02754v1 Announce Type: new Abstract: Personalization is a crucial capability of modern language agents.
By Peixuan Han, Hongyi Du, Jiayu Liu, Yihang Sun, Yutong Liu, Jiaxuan You
The paper introduces PRISM, a new framework for evaluating how well large language models (LLMs) maintain persona fidelity in dynamic dialogue. PRISM reframes the task as a structured inverse inference problem grounded in Systemic Functional Linguistics, breaking persona fidelity into Task Framing, Interpersonal Stance, and Linguistic Style dimensions. Experiments demonstrate that PRISM produces more accurate and stable judgments than existing holistic or static psychometric methods, offering a more reliable and auditable evaluation process.
By Mengfan Li, Zesheng Wei, Xuanhua Shi, Yang Deng