Dissociating the Internal Representations of Sycophancy in LLMs
arXiv:2607. 07003v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect.
arXiv:2607. 00415v1 Announce Type: cross Abstract: Authority bias poses a critical safety concern in language models: models systematically prioritize social cues from authority figures over factual consistency, swaying their answers based on source credibility rather than evidence.
arXiv:2607. 07003v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect.
arXiv:2608. 08026v1 Announce Type: new Abstract: We investigate how social authority (SA) signals interact with severity-based prioritization in large language models, operationalizing each axis as a model-elicited baseline -- the triage hierarchy and the SA hierarchy.
arXiv:2608. 15687v1 Announce Type: new Abstract: Sycophancy, the tendency of a language model to change its answer to match a user's stated belief, is a common alignment failure.
Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.
arXiv:2606. 05403v1 Announce Type: new Abstract: Language models increasingly act as epistemic proxies, synthesizing evidence from multiple sources to inform decisions.
arXiv:2606. 01637v1 Announce Type: cross Abstract: Large language models are increasingly used in multi-agent systems, where they see and respond to other agents' answers.
arXiv:2408. 05568v2 Announce Type: replace Abstract: Large Language Models (LLMs) exhibit potentially harmful biases that reinforce culturally embedded stereotypes, influence moral judgments, or amplify positive evaluations of majority groups.
arXiv:2606. 05976v1 Announce Type: new Abstract: Recent work shows that LLM agents struggle to correct errors in their own reasoning traces yet show markedly higher correction rates when identical claims appear under external sources.
arXiv:2605. 21006v2 Announce Type: replace Abstract: We study the effect of different persona on \textbf{sycophancy}: model's agreement with users even when the user is incorrect.
arXiv:2606. 05976v2 Announce Type: replace Abstract: Recent works show that LLM agents struggle to correct errors in their own reasoning traces, despite their ability to correct errors from external sources.
arXiv:2606. 11502v1 Announce Type: cross Abstract: Language models can state that "the Earth orbits the Sun" and, when role-playing Aristotle, assert the opposite.
arXiv:2603. 09986v3 Announce Type: replace-cross Abstract: Hallucinations, the tendency for large language models to provide responses with factually incorrect and unsupported claims, is a serious problem within natural language processing for which we do not yet have an effective solution to mitigate against.