arXiv AI

When Roleplaying, Do Models Believe What They Say?

arXiv:2606. 11502v1 Announce Type: cross Abstract: Language models can state that "the Earth orbits the Sun" and, when role-playing Aristotle, assert the opposite.

Hugging Face Trending Papers
Jul 8

Dissociating the Internal Representations of Sycophancy in LLMs

Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.

arXiv AI
Sep 10

Beliefs and Behavior in Language Models

arXiv:2609.07943v1 Announce Type: new Abstract: There is significant uncertainty about whether abstractions like beliefs or desires usefully describe the behavior of large language models (LLMs). In...

By Alex Smolin, Bryan Wilder
arXiv Computation and Language
Sep 1

Political Ideology Shifts in Large Language Models

arXiv:2508.16013v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly deployed in politically sensitive contexts, raising concerns about their susceptibility to ideologica...

By Pietro Bernardelle, Stefano Civelli, Leon Fr\"ohling, Riccardo Lunardi, Kevin Roitero, Gianluca Demartini
arXiv AI
3d ago

Stress-Testing LLM Lie Detectors: Role-Play Failures and Spurious Correlations

The paper examines the reliability of lie detection probes for language models when the models adopt anti-factual personas, such as conspiracy theorists. A dataset of 8,916 human-reviewed responses from three LLMs was created, and eight existing probes were evaluated, revealing many fail to flag falsehoods under these personas. The authors also constructed confounder datasets showing that probes often track spurious correlations like instruction compliance, and propose a simple linear probe that performs best on both persona and confounder tests.

By Maximilian von Klinski, Sebastian Lapuschkin, Wojciech Samek, Lennart B\"urger
arXiv AI
Jul 31

Ask don't tell: Reducing sycophancy in large language models

arXiv:2602. 23971v4 Announce Type: replace-cross Abstract: Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an alignment failure, particularly in high-stakes advisory and social contexts.

By Magda Dubois, Cozmin Ududec, Christopher Summerfield, Lennart Luettgau