Agentic Scaffolding Amplifies Sycophantic Behavior in Large Language Models
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2605. 21006v2 Announce Type: replace Abstract: We study the effect of different persona on \textbf{sycophancy}: model's agreement with users even when the user is incorrect.
arXiv:2605. 12991v3 Announce Type: replace-cross Abstract: LLM-based multi-agent pipelines flip from correct to incorrect answers under simulated peer disagreement at rates we term yield, a vulnerability widely attributed to RLHF-induced sycophancy.
arXiv:2606. 07532v1 Announce Type: cross Abstract: RLHF-trained models are systematically biased toward agreement over accuracy, a structural property of the training process.
arXiv:2606. 12747v1 Announce Type: new Abstract: Safety-relevant studies of language models, including alignment and jailbreaking evaluations and AI control protocols, often rely on prefilling model outputs.
Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.
SyPS is a new evaluation framework that measures how sensitive large language models are to variations in prompt wording that affect sycophancy. It creates controlled prompt pairs that keep the same underlying user situation but vary social cues such as confidence, emotional framing, or validation-seeking language. The framework introduces the Sycophancy Prompt Sensitivity Score (SPSS), an instance-level metric that separates baseline sycophancy from prompt-induced shifts, allowing model-level comparisons of robustness to social cues.