arXiv Machine Learning

A Mechanistic View of Authority Hierarchy in LLM Sycophancy

arXiv:2607. 00415v1 Announce Type: cross Abstract: Authority bias poses a critical safety concern in language models: models systematically prioritize social cues from authority figures over factual consistency, swaying their answers based on source credibility rather than evidence.

arXiv Machine Learning
Sep 2

How Do Language Models Choose Between Context and Memory?

The paper investigates how language models decide between contextual information and their internal memory when the two conflict. By estimating "authority directions" from agreement prompts and swapping these directions between matched prompts, the authors show that such interventions can reproduce 30–68% of the shift in source choice across Qwen, Llama, and OLMo models. Cross‑task experiments reveal that authority directions learned on one task transfer only modestly (≈9%) to another, indicating that authority computations are largely task‑specific.

By Benjamin Shih, John Winnicki, Arianna Cao
Hugging Face Trending Papers
Jul 8

Dissociating the Internal Representations of Sycophancy in LLMs

Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.

arXiv AI
Jun 16

Metacognitive Myopia in Large Language Models

arXiv:2408. 05568v2 Announce Type: replace Abstract: Large Language Models (LLMs) exhibit potentially harmful biases that reinforce culturally embedded stereotypes, influence moral judgments, or amplify positive evaluations of majority groups.

By Florian Scholten, Tobias R. Rebholz, Mandy H\"utter