Dissociating the Internal Representations of Sycophancy in LLMs
arXiv:2607. 07003v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect.
arXiv:2607. 04523v1 Announce Type: cross Abstract: Generic statements like "tigers are striped" and "cars have radios" communicate information that is, in general, true.
arXiv:2607. 07003v1 Announce Type: new Abstract: Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect.
arXiv:2501. 05844v4 Announce Type: replace Abstract: Causal Learning has emerged as a major theme of research in statistics and machine learning in recent years, promising computational techniques to reveal ``true'' causality.
Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.
arXiv:2606. 30815v1 Announce Type: cross Abstract: Recent work suggests that transformer language models show a bias towards human languages over unnatural ("impossible") languages argued to be unacquirable by humans.
When large language models (LLMs) fail to generalize or make haphazard errors in reasoning, it is often taken as evidence that LLMs are not truly reasoning, but rather performing a kind of pattern matching. The implication is that people's behavior does not exhibit the same types of failures because human reasoning uses principled and abstract world models.
The paper investigates the relationship between a large language model’s internal probability distribution and its verbalized confidence statements. By systematically manipulating training and in‑context data, the authors show that both internal and verbalized probabilities are influenced by distributional and asserted uncertainty in the data. They find that verbalized probabilities align with internal ones beyond what would be expected if they tracked the same sources independently, indicating that verbalized confidence can serve as a probe of the model’s internal distribution.
arXiv:2609.16854v1 Announce Type: new Abstract: Probability is fundamental to theories of language comprehension, production, acquisition, and evolution, as well as to large language models. Existing...
arXiv:2606. 13607v1 Announce Type: new Abstract: When large language models (LLMs) fail to generalize or make haphazard errors in reasoning, it is often taken as evidence that LLMs are not truly reasoning, but rather performing a kind of pattern matching.
arXiv:2608.28924v1 Announce Type: new Abstract: Linguistic theory has long recognized cross-linguistic syntactic regularities, leading to claims that these similar structures are processed by similar...
arXiv:2609.07943v1 Announce Type: new Abstract: There is significant uncertainty about whether abstractions like beliefs or desires usefully describe the behavior of large language models (LLMs). In...
arXiv:2609.22409v1 Announce Type: new Abstract: Understanding contextual causality is critical for large language models (LLMs), as it enables them to accurately identify causal relations in specific...
arXiv:2604. 14180v2 Announce Type: replace-cross Abstract: We train a 318M-parameter Transformer language model from scratch on a curated corpus of 1.