arXiv AI
Aug 28

Do LLMs Understand Personality? Rethinking Persona Fidelity Evaluation through Structured Behavioral Inference

The paper introduces PRISM, a new framework for evaluating how well large language models (LLMs) maintain persona fidelity in dynamic dialogue. PRISM reframes the task as a structured inverse inference problem grounded in Systemic Functional Linguistics, breaking persona fidelity into Task Framing, Interpersonal Stance, and Linguistic Style dimensions. Experiments demonstrate that PRISM produces more accurate and stable judgments than existing holistic or static psychometric methods, offering a more reliable and auditable evaluation process.

By Mengfan Li, Zesheng Wei, Xuanhua Shi, Yang Deng
arXiv AI
6d ago

Do Personality-Tuned LLMs Make Better Social Agents?

The paper examines whether fine‑tuning large language models (LLMs) with personality‑labelled data improves their ability to act as socially interactive agents. Two small open‑weight LLMs were fine‑tuned on a corpus of personality‑labelled social media posts and dialogues, and the resulting models were evaluated in various social interaction scenarios by independent LLM judges. The findings show that the fine‑tuned models do not outperform their baseline counterparts in role‑playing personalities, though they offer comparable text quality and increased linguistic diversity for the Qwen models; low inter‑rater agreement limits confidence in the results, suggesting future work should focus on training data quality and domain alignment.

By Tim Krabbe, Xiaodan Shi
arXiv AI
Jun 18

How Well Do Large Language Models Capture Human Personality?

arXiv:2606. 18263v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to simulate human populations via persona prompting, often under the assumptions that richer persona descriptions improve behavioral fidelity, similarly sized attribute combinations are equally simulatable, and persona definitions generalize across tasks.

By Aanisha Bhattacharyya, Yaman Kumar Singla, Rajiv Ratn Shah, Changyou Chen, Jitendra Ajmera
arXiv AI
Sep 16

RoleBreak: Benchmarking Long-Horizon Role-Playing Robustness in Spoken Dialogue

RoleBreak is an open benchmark designed to evaluate long‑horizon role‑playing robustness in spoken dialogue systems. It includes 310 character‑based and user‑centered roles, 6,688 human‑verified dialogue turns, and 11,743 fine‑grained evaluation criteria, with 1,856 turns specifically targeting expressive vocal emotion. The benchmark stresses role consistency, interaction quality, safety, and affect over extended conversations, and the authors evaluated nine system configurations across full‑duplex, omni‑modal, and cascaded ASR–LLM–TTS paradigms.

By Yuqi Wang, Fengyuan Liu, Haochen Luo, Zhiqi Yu, Qi Liu