arXiv AI
Jun 2

RoleCDE:Benchmarking and Mitigating Role-Alignment Trade-offs in Role-Playing Agents

arXiv:2606. 01552v1 Announce Type: new Abstract: Role-playing agents(RPAs) are widely used to steer large language models(LLMs) toward role-consistent behavior, yet existing benchmarks mainly evaluate surface-level fidelity and offer limited insight into decision making under role-alignment value conflicts.

By Huayi Lai, Shichao Song, Simin Niu, Hanyu Wang, Jiawei Yang, Zhouxing Wang, Zhiqiang Yin, Xun Liang
arXiv Computation and Language
Sep 1

Identifying and Mitigating Bottlenecks in Role-Playing Agents: A Systematic Study of Disentangling Character Profile Axes

The paper introduces a diagnostic framework to disentangle the effects of character profile axes—Familiarity, Structure, and Disposition—on large language model role‑playing agents. Experiments on 211 personas and five LLMs show that Familiarity and Structure have little impact, whereas Disposition, particularly immoral traits, consistently degrades performance. The authors propose Field‑Aware Contrastive Decoding (FACD), a training‑free method that mitigates this performance gap without harming moral‑character performance.

By Yonghyun Jun, Junhyuk Choi, Jeonghyun Park, Jihyeong Park, Liu Nicole Geumheon, Hwanhee Lee