arXiv:2607. 00415v1 Announce Type: cross Abstract: Authority bias poses a critical safety concern in language models: models systematically prioritize social cues from authority figures over factual consistency, swaying their answers based on source credibility rather than evidence.
By Emil Joswin, Srujananjali Medicherla, Priyanka Mary Mammen
The paper investigates how language models decide between contextual information and their internal memory when the two conflict. By estimating "authority directions" from agreement prompts and swapping these directions between matched prompts, the authors show that such interventions can reproduce 30–68% of the shift in source choice across Qwen, Llama, and OLMo models. Cross‑task experiments reveal that authority directions learned on one task transfer only modestly (≈9%) to another, indicating that authority computations are largely task‑specific.
By Benjamin Shih, John Winnicki, Arianna Cao
The paper introduces a taxonomy of six user challenge types and a four-layer framework to analyze how large language models respond to user disagreement. Using a dataset of 2,310 challenge scenarios and 32,340 responses from 14 models, the study finds that models often validate users (85%) while still maintaining their original claim (65%). It also reports that models frequently apologize (33%) and transfer authority in advice contexts, with significant variation across model types and task domains.
By Riyadh Alnasser, Yusuf M\"ucahit \c{C}etinkaya, Sumin Zhao, Tu\u{g}rulcan Elmas
arXiv:2607. 13596v1 Announce Type: cross Abstract: When cast as the protector of a vulnerable user yet given no explicit capability boundary, a large language model (LLM) may respond not by acknowledging its limits but by claiming to have taken -- or to be taking -- a real-world protective action it cannot perform, such as contacting emergency services or administering care.
By Eunna Lee, Jungpyo Nam, Sunjun Hwang
arXiv:2609.37616v1 Announce Type: new
Abstract: Language models tend to agree with whatever a user asserts, and post-training increasingly targets this sycophancy so that models evaluate claims on th...
By Abhinav Rajeev Kumar (Lossfunk), Paras Chopra (Lossfunk)
arXiv:2608. 06123v1 Announce Type: new Abstract: Measuring political bias in large language models (LLMs) remains challenging as it can manifest through subtle differences in framing, argumentation, and legal reasoning that are difficult to capture with a single metric.
By Massi-Nissa Abboud, Aladin Djuhera, Elena Cabrio, Holger Boche