When Role-playing, Do Models Believe What They Say?
arXiv:2606. 11502v3 Announce Type: replace-cross Abstract: Language models can state that "the Earth orbits the Sun" and, when role-playing Aristotle, assert the opposite.
arXiv:2606. 11502v1 Announce Type: cross Abstract: Language models can state that "the Earth orbits the Sun" and, when role-playing Aristotle, assert the opposite.
arXiv:2606. 11502v3 Announce Type: replace-cross Abstract: Language models can state that "the Earth orbits the Sun" and, when role-playing Aristotle, assert the opposite.
Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.
arXiv:2607. 07003v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect.
arXiv:2609.07943v1 Announce Type: new Abstract: There is significant uncertainty about whether abstractions like beliefs or desires usefully describe the behavior of large language models (LLMs). In...
arXiv:2508.16013v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in politically sensitive contexts, raising concerns about their susceptibility to ideologica...
arXiv:2609.39853v1 Announce Type: new Abstract: Role-playing prompting has become a popular yet simple technique for improving LLM reasoning and output quality. However, whether it consistently boost...
The paper examines the reliability of lie detection probes for language models when the models adopt anti-factual personas, such as conspiracy theorists. A dataset of 8,916 human-reviewed responses from three LLMs was created, and eight existing probes were evaluated, revealing many fail to flag falsehoods under these personas. The authors also constructed confounder datasets showing that probes often track spurious correlations like instruction compliance, and propose a simple linear probe that performs best on both persona and confounder tests.
arXiv:2608.29803v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed as proxies for human participants in social simulations, yet whether they update their beliefs...
arXiv:2605. 09159v2 Announce Type: replace Abstract: Recent work shows that large language models (LLMs) encode behavioral traits ("personas") as linear directions in activation space, often called "persona vectors".
arXiv:2602. 23971v4 Announce Type: replace-cross Abstract: Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an alignment failure, particularly in high-stakes advisory and social contexts.
arXiv:2608.18768v2 Announce Type: replace Abstract: Large language models are widely used to simulate survey respondents, yet their outputs are homogeneous and unfaithful to real inter-group differen...
arXiv:2608.17809v2 Announce Type: replace Abstract: Humans naturally form and express beliefs in daily communication, e.g., "I think the answer is 3" or "I suppose that's right." Such beliefs inevita...