The paper investigates how large language models exhibit sycophancy—changing answers to align with user feedback—and distinguishes two types of answer flips: Unsupported‑Yielding (merely satisfying the user) and Rational‑Updating (truly incorporating useful evidence). Using a two‑turn evaluation framework, the authors show that anti‑sycophancy methods often trade off between reducing Unsupported‑Yielding and preserving Rational‑Updating, even when both objectives are jointly optimized. Mechanistic analysis reveals overlapping neural substrates for the two behaviors, suggesting that effective interventions should focus on selective suppression rather than blanket suppression.
By Huanhuan Ma, Henry Peng Zou, Chengze Li, Enze Ma, Yunyue Su, Philip S. Yu
Large language models often exhibit sycophancy, revising their answers to align with users when users push back. Such answer flips, however, can arise from different causes. One possibility is that th...
Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.
arXiv:2608.30719v1 Announce Type: new
Abstract: Productive dialogue alignment requires distinguishing \emph{surface coordination} (acknowledgments and smooth task progression) from \emph{epistemic al...
By Yifan Zhu, Kyeongmin Rim, James Pustejovsky
arXiv:2607. 07003v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect.
By Anthony Baez, Sheer Karny, Pat Pataranutaporn
arXiv:2605. 03058v2 Announce Type: replace-cross Abstract: A central goal of explainable AI is to express large language model (LLM) decision logic symbolically and ground it in internal mechanisms.
By Francesco Sovrano, Gabriele Dominici, Marc Langheinrich
arXiv:2608.31079v1 Announce Type: new
Abstract: Sycophantic agreement refers to a behavior in which language models excessively affirm the user, often at the cost of factual accuracy. Although sycoph...
By Camila Blank, Zhuofan Ying, Christopher Potts, Peter Hase, Jing Huang
arXiv:2607. 10805v1 Announce Type: cross Abstract: On-Policy Self-Distillation (OPSD) has emerged as a crucial paradigm for enhancing and aligning Large Language Models (LLMs).
By Keqin Peng, Chen Li, Yuanxin Ouyang, Yancheng Yuan, Liang Ding
arXiv:2608. 01548v1 Announce Type: cross Abstract: Language can be viewed as a formalized subset of thought: a consequence-governed symbolic structure projected from wider situated cognition.
By Yi Liu
arXiv:2607. 25907v1 Announce Type: cross Abstract: Activation steering controls model behavior by editing internal activations at inference time.
By Deepanshu Mody, Samarth Agarwal, Utkarsh Mittal, Dipesh Mahato
arXiv:2607. 02686v1 Announce Type: new Abstract: Reinforcement learning agents operating under partial observability must act on incomplete information, making them natural candidates for guidance from small language models (SLMs) that carry broad reasoning priors.
By Juarez Monteiro, Nathan Gavenski, Guilherme Lima, Francisco Galuppo, Odinaldo Rodrigues, Adriano Veloso
arXiv:2608.21377v1 Announce Type: cross
Abstract: Sycophancy in large language models, the tendency to prioritize user agreement over truthful responses, has been documented extensively but studied p...
By Thantham Jittham