arXiv Computation and Language By Jongkyung Shin, Inkyu Lee, Chiehyeon Lim

Controlling and Assessing Appropriate Persona Use in LLM-based Dialogue Generation

Read the original on arXiv Computation and Language →

The paper examines how large language models (LLMs) tend to overuse persona attributes in persona-based dialogue generation, producing unnatural responses. It identifies a systematic bias in LLMs to incorporate all provided persona details and shows that current metrics cannot assess contextual appropriateness. To address this, the authors introduce Self-CONtrastive Persona Overuse Suppression (SCONPOS), which intervenes in the prompt encoding stage to reduce overuse, and propose the Persona Appropriateness Score (PAS), a new metric that penalizes both overuse and underuse of persona attributes.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Aug 12

Do LLMs Benefit From Their Own Words?

arXiv:2602. 24287v2 Announce Type: replace-cross Abstract: In multi-turn conversations, large language models typically condition on the full conversation history: both past user prompts and assistant responses.

By Jenny Y. Huang, Leshem Choshen, Wei Sun, Omar Khattab, Ram\'on Fernandez Astudillo, Mehul Damani, Tamara Broderick, Jacob Andreas
arXiv AI
Aug 28

Do LLMs Understand Personality? Rethinking Persona Fidelity Evaluation through Structured Behavioral Inference

The paper introduces PRISM, a new framework for evaluating how well large language models (LLMs) maintain persona fidelity in dynamic dialogue. PRISM reframes the task as a structured inverse inference problem grounded in Systemic Functional Linguistics, breaking persona fidelity into Task Framing, Interpersonal Stance, and Linguistic Style dimensions. Experiments demonstrate that PRISM produces more accurate and stable judgments than existing holistic or static psychometric methods, offering a more reliable and auditable evaluation process.

By Mengfan Li, Zesheng Wei, Xuanhua Shi, Yang Deng