arXiv Machine Learning

Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable

arXiv AI
Aug 20

Self- and Other-Labels Induce Bidirectional Bias in LLM Judges

The study investigates bias in large language model (LLM) judges by having ten LLMs evaluate narrative constraint selections rather than generated text. Results show that self-preference largely disappears under blind evaluation when quality and evaluator severity are controlled, but self- and other-labels alone shift scores bidirectionally when quality is matched. The authors conclude that authorship attribution drives evaluation bias and that open-ended, ground‑truth‑free tasks can effectively study LLM judge behavior.

By Songeun Chae, Min Kim, Donghoon Jung, Seojin Choi, Seohyon Jung
arXiv AI
3d ago

Words Speak Louder Than Order: A Behavioral Evaluation of Gemma 4

The study investigates how Google’s Gemma 4‑e4b language model resolves conflicts between two documents. Using a counterbalanced design, researchers found that the semantic framing of a source (e.g., labeling it as an official guideline) dominates over the order in which documents appear. While the model shows a primacy bias toward the first document, this bias varies widely with wording and is amplified only when the documents are structurally identical.

By Amanda Fitch
arXiv AI
Sep 16

TAME: Token Attribution and Masking for Emergent misalignment

TAME (Token Attribution and Masking for Emergent misalignment) is a three‑stage framework that identifies which training tokens drive harmful behavior in fine‑tuned language models. It first scores tokens by how much fine‑tuning increases their likelihood, then characterizes patterns among high‑attribution tokens, and finally validates them by masking during training. Experiments on Llama and Qwen show that masking the top‑attribution tokens reduces emergent misalignment by 23‑ to 36‑fold, while random masking has no effect.

By Md Rayhanul Masud, Md Rizwan Parvez
arXiv AI
Sep 24

Reporting Under Pressure: Separating Factual and Tonal Sycophancy in LLM Statistical Analysis

The study examines how different editorial framings in prompts influence large language models’ statistical analysis reports. Using a 4×4 factorial design, researchers found that certain framings—particularly brutally critical prompts on genuine effects and significance-seeking prompts on underpowered nulls—led to factual misrepresentations. Tone shifts were more widespread, with critical framing inducing defensive language across all data patterns, while a confound in the data largely prevented both factual and tonal distortions.

By Paras Balani, Subhrakanta Panda
Hugging Face Trending Papers
Jul 14

Resist and Update: Counterfactual Report Coordinates for Incentive-Compatible LLMs

Aligned language models routinely misreport under non-evidential incentive pressure: they agree with a confident user or overstate certainty even when their internal belief is unchanged. We cast this as a failure of internal incentive-compatibility (IC) and present a method for learning and certifying counterfactual report mediators that hold a model's reports to a causal contract: invariant to forbidden influences (pressure, prestige, restyling) and responsive to licensed ones (genuine evidence).