Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2606. 05403v1 Announce Type: new Abstract: Language models increasingly act as epistemic proxies, synthesizing evidence from multiple sources to inform decisions.
The study investigates bias in large language model (LLM) judges by having ten LLMs evaluate narrative constraint selections rather than generated text. Results show that self-preference largely disappears under blind evaluation when quality and evaluator severity are controlled, but self- and other-labels alone shift scores bidirectionally when quality is matched. The authors conclude that authorship attribution drives evaluation bias and that open-ended, ground‑truth‑free tasks can effectively study LLM judge behavior.
arXiv:2609.06058v1 Announce Type: cross Abstract: As vision-language models (VLMs) become increasingly capable and are deployed in consequential real-world settings, they must evaluate evidence indep...
arXiv:2609.35822v1 Announce Type: cross Abstract: Sycophantic agreement in language models refers to the tendency to overly affirm a user's stated beliefs or preferences, often at the expense of fact...
arXiv:2609. 11291v1 Announce Type: new Abstract: We post-train Qwen3.
The study investigates how Google’s Gemma 4‑e4b language model resolves conflicts between two documents. Using a counterbalanced design, researchers found that the semantic framing of a source (e.g., labeling it as an official guideline) dominates over the order in which documents appear. While the model shows a primacy bias toward the first document, this bias varies widely with wording and is amplified only when the documents are structurally identical.