arXiv Computation and Language

Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It

Hugging Face Trending Papers
Aug 18

Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It

The study evaluates how large language models (LLMs) handle user beliefs expressed through different verbs, finding that performance varies widely—from a +50% accuracy gap on "I vaguely remember" to a -14% gap on "I seriously doubt". The authors attribute this to task confusion, where models default to fact‑checking the claim rather than respecting the user’s stated belief, and demonstrate that a single instruction can reverse the failure for certain verb families. Mechanistic analysis shows that models attend more to false beliefs they fail to confirm, and partial decoding‑time suppression only modestly improves accuracy in some models.

Hugging Face Trending Papers
Jul 8

Dissociating the Internal Representations of Sycophancy in LLMs

Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.

arXiv AI
Jul 7

The Anatomy of Uncertainty in LLMs

arXiv:2603. 24967v2 Announce Type: replace Abstract: Understanding why a large language model (LLM) is uncertain about the response is important for their reliable deployment.

By Aditya Taparia, Ransalu Senanayake, Kowshik Thopalli, Vivek Narayanaswamy