arXiv AI

Point of Order: Action-Aware LLM Persona Modeling for Data-Grounded Civic Deliberation

arXiv:2511. 17813v3 Announce Type: replace-cross Abstract: LLM-based simulations can enable controlled studies of civic deliberation, but current systems lack speaker-attributed data and methods for evaluating long-form institutional behavior.

arXiv AI
Sep 10

From Simulated Citizens to Simulated Deliberation: Challenges in Representation and Interaction

The paper investigates whether large language model (LLM) agents can simulate public deliberation by reflecting population opinion patterns and producing interaction-driven opinion change. Using census‑grounded Korean personas debating real policy questions, the study finds that persona agents fail to reliably reproduce population opinion patterns, often concentrating responses and reversing demographic differences. While deliberations generate reasoned, reciprocal arguments and some stance movement, much of this change occurs without peer exchange, and anchoring agents to population‑informed starting positions suppresses updating, indicating that population representation, argument generation, and interaction‑driven opinion change do not necessarily align.

By Chaemin Jang, Junsik Min, Jaewoo Choi, Donggyu Lee, Haiin Lee, Junyoung Park, Namhee Kim, Hyunwoo Kim, Jungwon Kim, Juho Kim, Nuri Kim, Jihee Kim
arXiv AI
Sep 4

Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting

The paper introduces CAPA, a Collaborative Agent Predictive Architecture designed to give large language model (LLM) agents situational awareness in online meetings. CAPA uses a Perceiver to update meeting state, a Predictor to forecast conversation flow, a Controller to decide speaking actions, and a Generator to phrase contributions. Evaluated on 137 AMI meetings, CAPA reduces the silence rate from 51.4% to 2.5%, doubles credited recovery, and maintains low hallucination, demonstrating that structured state tracking is key to effective delegation.

By Muneeb Khan, Frederic Kirstein, Terry Ruas, Bela Gipp
arXiv Computation and Language
Sep 11

Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction

The paper introduces CoCoEval, a framework for evaluating large language model (LLM)–simulated conversations by detecting 10 types of inconsistent and uncollaborative behaviors at the turn level. Using CoCoEval, the authors compare human conversations with those generated by GPT‑4.1, GPT‑5.1, and Claude Opus 4, finding that LLMs produce far fewer such behaviors under vanilla prompting and that prompt engineering or fine‑tuning often over‑produces specific behaviors. The study highlights gaps between human and LLM‑simulated interactions that conventional Likert‑scale evaluations miss, raising concerns about using LLMs as proxies for human social interaction.

By Ryo Kamoi, Ameya Godbole, Binglin Zhou, Xiaoxin Lu, Longqi Yang, Rui Zhang, Mengting Wan, Pei Zhou
arXiv AI
Aug 28

Do LLMs Understand Personality? Rethinking Persona Fidelity Evaluation through Structured Behavioral Inference

The paper introduces PRISM, a new framework for evaluating how well large language models (LLMs) maintain persona fidelity in dynamic dialogue. PRISM reframes the task as a structured inverse inference problem grounded in Systemic Functional Linguistics, breaking persona fidelity into Task Framing, Interpersonal Stance, and Linguistic Style dimensions. Experiments demonstrate that PRISM produces more accurate and stable judgments than existing holistic or static psychometric methods, offering a more reliable and auditable evaluation process.

By Mengfan Li, Zesheng Wei, Xuanhua Shi, Yang Deng
Hugging Face Trending Papers
Sep 3

Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting

The paper introduces CAPA, a Collaborative Agent Predictive Architecture designed to improve large language model (LLM) participation in online meetings. CAPA tracks meeting state with a Perceiver, predicts conversation flow, decides when and what to speak, and generates contributions in the participant’s style, all while being calibrated by judges. In experiments on 137 AMI meetings, CAPA cuts the LLM’s silence rate from 51.4% to 2.5%, doubles credited recovery, and maintains low hallucination.

arXiv AI
Sep 10

A Three-Tier Persona Vector for Controllable User Simulation in Agentic Evaluation

The paper introduces a three-tier persona vector for user simulation in evaluating LLM agents, comprising 23 dimensions across demographics, behavioral traits, and emotional states, plus a query-complexity overlay. It demonstrates that these nuanced personas generate diverse, scenario-reactive conversations, leading to significant variations in agent goal achievement and compliance across different contexts. The model’s design allows for reproducible, auditable user behavior patterns without relying on learned covariance matrices.

By Rahul Khedar, Eshita, Sneha Teja Sree Reddy Thondapu, Mayank Malhotra, Arup Kumar Das, Jitesh Chandra Mishra, Arun Menon, Avinash Karn, Mouli V
arXiv AI
Jun 16

State-Grounded Multi-Agent Synthetic Data Generation for Tool-Augmented LLMs

arXiv:2606. 16307v1 Announce Type: new Abstract: Training tool-augmented LLM agents requires large corpora of multi-turn, tool-grounded conversational data that is expensive to annotate, privacy-constrained in production settings, and largely absent from public datasets.

By Rahul Khedar, Eshita, Sneha Teja Sree Reddy Thondapu, Mayank Malhotra, Arup Das, Jitesh Chandra, Yun-Shiuan Chuang, Chaitanya Kulkarni, Arun Menon, Linsey Pang, Avinash Karn, Mouli V, Prakhar Mehrotra