Dissociating the Internal Representations of Sycophancy in LLMs
arXiv:2607. 07003v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect.
arXiv:2607. 07003v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect.
Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.
arXiv:2602. 23971v4 Announce Type: replace-cross Abstract: Sycophancy, the tendency of large language models to favour user-affirming responses over critical engagement, has been identified as an alignment failure, particularly in high-stakes advisory and social contexts.
arXiv:2609.37616v1 Announce Type: new Abstract: Language models tend to agree with whatever a user asserts, and post-training increasingly targets this sycophancy so that models evaluate claims on th...
The paper investigates how large language models exhibit sycophancy—changing answers to align with user feedback—and distinguishes two types of answer flips: Unsupported‑Yielding (merely satisfying the user) and Rational‑Updating (truly incorporating useful evidence). Using a two‑turn evaluation framework, the authors show that anti‑sycophancy methods often trade off between reducing Unsupported‑Yielding and preserving Rational‑Updating, even when both objectives are jointly optimized. Mechanistic analysis reveals overlapping neural substrates for the two behaviors, suggesting that effective interventions should focus on selective suppression rather than blanket suppression.
arXiv:2509.21305v4 Announce Type: replace Abstract: Large language models (LLMs) often exhibit sycophantic behaviors -- such as excessive agreement with or flattery of the user -- but it is unclear w...
Large language models often exhibit sycophancy, revising their answers to align with users when users push back. Such answer flips, however, can arise from different causes. One possibility is that th...
arXiv:2608. 15687v1 Announce Type: new Abstract: Sycophancy, the tendency of a language model to change its answer to match a user's stated belief, is a common alignment failure.
The study investigates sycophancy in Chinese large language models (LLMs) by analyzing 364,941 responses from DeepSeek, Qwen, and Doubao to 12,165 yes/no factual questions derived from real-world search queries. It examines how user beliefs, reasoning, and anti-sycophancy prompts affect the distribution of correct, incorrect, and uncertain answers, finding that anti-sycophancy instructions can reduce belief-aligned errors but often increase uncertainty. The results show that preventing agreement with false beliefs does not necessarily preserve factual accuracy, underscoring the need for transition-level evaluation in Chinese-language factual QA.
arXiv:2609.26579v1 Announce Type: new Abstract: A central concern with language models is sycophancy: their tendency to defer to users' views at the expense of independent substantive judgment. In pa...
arXiv:2609.37491v1 Announce Type: cross Abstract: Retrieval-augmented language models are expected to answer from the retrieved evidence, but in practice they often keep answering when that evidence...
arXiv:2603. 01437v2 Announce Type: replace Abstract: As chain of thought (CoT) has become central to scaling reasoning capabilities in large language models (LLMs), it has also emerged as a promising tool for interpretability, suggesting the opportunity to understand model decisions through verbalized reasoning.