arXiv:2609.22607v1 Announce Type: new
Abstract: We argue here that the current dominant practice in LLM human simulation: prompting instruction-tuned assistant language models to role-play personas,...
By Minwoo Kang, T\'ea Wright, Seun Eisape, Ayush Raj, Suhong Moon, Joseph Suh, Alane Suhr, David M. Chan, John Canny
arXiv:2609.26780v1 Announce Type: cross
Abstract: Long-term conversational memory in multi-party settings requires more than retrieving relevant content from long-term conversations: it must distingu...
By Haobo Zheng, Tan Tang, Yan Chen, Weijie Wang, Yingcai Wu
arXiv:2609.18935v1 Announce Type: new
Abstract: A game character should not have to reread its entire life before every conversation. For locally deployed language-model characters, however, revising...
By Zimu Xu
DiaRelay introduces a lightweight adapter that lets large language models maintain a constant‑size dialogue‑level memory for emotion recognition in conversation. It builds on LoRA by adding a Selective Relay Memory Transition that aggregates useful historical evidence into a bounded memory, and a Dual‑axis Relay Memory Read that uses this memory to modulate low‑rank feature transformations. Experiments show DiaRelay achieves state‑of‑the‑art weighted F1 and accuracy on MELD with only 7.1 M additional trainable parameters, while also performing competitively on IEMOCAP.
By Zihao Zhou, Bin Yang, Jinghui Qin, Kebing Jin
arXiv:2607. 17191v1 Announce Type: new Abstract: Human-like private chat requires more than fluent response generation: a system must preserve persona, relationship, memory, bounded knowledge, medium-specific timing, and a coherent multi-turn arc.
By Wentao Liu, Siyu Song, Xi Chen, Youjia Li, Xiaokun Wang, Min Ji, Ji Wang
RoleBreak is an open benchmark designed to evaluate long‑horizon role‑playing robustness in spoken dialogue systems. It includes 310 character‑based and user‑centered roles, 6,688 human‑verified dialogue turns, and 11,743 fine‑grained evaluation criteria, with 1,856 turns specifically targeting expressive vocal emotion. The benchmark stresses role consistency, interaction quality, safety, and affect over extended conversations, and the authors evaluated nine system configurations across full‑duplex, omni‑modal, and cascaded ASR–LLM–TTS paradigms.
By Yuqi Wang, Fengyuan Liu, Haochen Luo, Zhiqi Yu, Qi Liu
A game character should not have to reread its entire life before every conversation. For locally deployed language-model characters, however, revising a few memories can invalidate a long reusable pr...
arXiv:2605. 09159v2 Announce Type: replace Abstract: Recent work shows that large language models (LLMs) encode behavioral traits ("personas") as linear directions in activation space, often called "persona vectors".
By Nils A. Herrmann, Leander Girrbach, Kirill Bykov, Zeynep Akata
arXiv:2510. 16392v3 Announce Type: replace Abstract: Personalized and continuous interactions are critical for LLM-based conversational agents, yet finite context windows and static parametric memory hinder the modeling of long-term, cross-session user states.
By Ao Tian, Yunfeng Lu, Xinxin Fan, Changhao Wang, Lanzhi Zhou, Yeyao Zhang, Yanfang Liu
arXiv:2606. 05553v1 Announce Type: cross Abstract: Role-playing language agents (RPLAs) should play characters whose values and behavior evolve as the story progresses, not maintain a fixed persona.
By Woojung Song, Nalim Kim, Sangjun Song, Chaewon Heo, Jongwon Lim, Yohan Jo
VoiceLongMemEval (VLME) is a new benchmark that tests AI assistants on their ability to remember how users sounded by incorporating paralinguistic metadata—such as emotion labels, prosody descriptors, and voice events—into each conversational turn. The benchmark uses a three‑stage adversarial gate to ensure that models cannot succeed with transcript alone, revealing a significant affect gap: models gain 0.09 to 0.38 accuracy when provided with paralinguistic cues, and audio‑native models outperform standard ASR pipelines in extracting these signals. The dataset and code will be released upon acceptance.
By Ramit Pahwa, Parivesh Priye, Apoorva Beedu
RENDER is a benchmark that controls the reader‑facing artifact in memory and RAG evaluations while keeping the conversation fixed. It introduces a five‑level packet ladder and deterministic templates that mimic ChatGPT‑style entries, LangChain summaries, MemGPT‑style typed records, and raw conversation. Experiments on 500 LongMemEval questions across nine models show that matched‑budget packets outperform raw dialogue by 42.4–72.6 points, and that ChatGPT‑style entries often score higher than raw conversation, with effects persisting under retrieval noise and transferring to HotpotQA.
By Yuan Si, Simeng Han, Daming Li, Jialu Zhang