Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.
arXiv:2509.21305v4 Announce Type: replace
Abstract: Large language models (LLMs) often exhibit sycophantic behaviors -- such as excessive agreement with or flattery of the user -- but it is unclear w...
By Daniel Vennemeyer, Phan Anh Duong, Tiffany Zhan, Tianyu Jiang
arXiv:2609.35822v1 Announce Type: cross
Abstract: Sycophantic agreement in language models refers to the tendency to overly affirm a user's stated beliefs or preferences, often at the expense of fact...
By Sixing Chen, Zhuofan Josh Ying, Logan Riggs Smith, Jeremy Wertheimer, Natalie Shapira
arXiv:2604.03058v3 Announce Type: replace-cross
Abstract: LLMs can be socially sycophantic, affirming users when they ask questions like "am I in the wrong?" rather than providing genuine assessment....
By Myra Cheng, Isabel Sieh, Humishka Zope, Sunny Yu, Lujain Ibrahim, Aryaman Arora, Jared Moore, Desmond Ong, Dan Jurafsky, Diyi Yang
arXiv:2608.29198v1 Announce Type: new
Abstract: As Large Language Models (LLMs) increasingly encourage users to disclose personal profiles for tailored assistance, measuring their political alignment...
By Li-Ni Fu, Chang-Chih Meng, Chien-Hua Chen, Hen-Hsen Huang, I-Chen Wu
arXiv:2606. 08076v1 Announce Type: cross Abstract: Large Language Models (LLMs) can generate high-quality arguments, yet their ability to engage in nuanced and persuasive communicative actions remains largely unexplored.
By Esra D\"onmez, Agnieszka Falenska
arXiv:2608.17809v2 Announce Type: replace
Abstract: Humans naturally form and express beliefs in daily communication, e.g., "I think the answer is 3" or "I suppose that's right." Such beliefs inevita...
By Quang Minh Nguyen, Luis Frentzen Salim
arXiv:2609.07943v1 Announce Type: new
Abstract: There is significant uncertainty about whether abstractions like beliefs or desires usefully describe the behavior of large language models (LLMs). In...
By Alex Smolin, Bryan Wilder
arXiv:2609.26579v1 Announce Type: new
Abstract: A central concern with language models is sycophancy: their tendency to defer to users' views at the expense of independent substantive judgment. In pa...
By Calvin Isley, Johann Gaebler, Max Lamparth, Julia Minson, Sharad Goel
arXiv:2606. 11502v1 Announce Type: cross Abstract: Language models can state that "the Earth orbits the Sun" and, when role-playing Aristotle, assert the opposite.
By Benjamin Sturgeon, David Africa, Sid Black
arXiv:2605.30381v2 Announce Type: replace-cross
Abstract: When a language model is fine-tuned to produce systematically incorrect responses, does this training leave a structured, linearly recoverabl...
By Vahideh Zolfaghari
arXiv:2606. 11205v1 Announce Type: cross Abstract: Activation steering can shift LLM behaviour, but standard evaluations do not typically test whether a sycophancy-reduction direction also suppresses agreement with factually correct statements.
By Matthew James Buchan