arXiv AI

Query Timing Produces Opposite Positional Biases Between LLMs and Humans

arXiv:2608. 12387v1 Announce Type: cross Abstract: Positional biases such as recency and primacy effects have been documented in large language models (LLMs), yet the underlying mechanism by which these models make their evaluations remains poorly understood.

arXiv AI
6d ago

Words Speak Louder Than Order: A Behavioral Evaluation of Gemma 4

The study investigates how Google’s Gemma 4‑e4b language model resolves conflicts between two documents. Using a counterbalanced design, researchers found that the semantic framing of a source (e.g., labeling it as an official guideline) dominates over the order in which documents appear. While the model shows a primacy bias toward the first document, this bias varies widely with wording and is amplified only when the documents are structurally identical.

By Amanda Fitch
Hugging Face Trending Papers
Jul 8

Dissociating the Internal Representations of Sycophancy in LLMs

Large Language Models (LLMs) frequently exhibit sycophancy, where they agree with a user's statement even when incorrect. While sycophancy is often treated as a single defined behavior, it can manifest in substantially distinct ways and circumstances, raising the question of whether this multi-faceted nature is reflected in its internal mechanisms.

arXiv Machine Learning
Sep 11

When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text

Large language models (LLMs) are increasingly used to assess social bias in text, but the passages they evaluate often contain surface noise such as typos and broken punctuation. This study applied five realistic noise conditions at varying intensities to 3,822 stereotype‑related responses and compared bias judgments on noisy versus original text. The findings show that noise disproportionately turns neutral judgments into biased ones—up to 120 times more likely—while rarely converting biased judgments into neutral ones, and that the most fragile LLM judge exhibits the greatest distortion at mild noise levels. As LLMs become more robust, the bias distortion tends toward parity rather than reversal, meaning bias measured on noisy text is systematically overestimated, especially in fairness‑critical categories.

By DongHyun Ryu, Jaehyeok Lee, YeongJun Hwang, JinYeong Bak
arXiv Computation and Language
Sep 1

Different Demographic Cues Yield Inconsistent Conclusions About LLM Personalization and Bias

The paper examines how large language models (LLMs) respond to different demographic cues—such as names—when users seek advice, focusing on race and gender in a U.S. context. It finds that using different cues for the same group leads to only partially overlapping changes in model responses, producing inconsistent conclusions about personalization and unstable bias metrics. The authors argue that LLMs react to linguistic signals tied to cues rather than to stable demographic categories, and they call for evaluations that use multiple cues and consider underlying mechanisms.

By Manuel Tonneau, Neil K. R. Sehgal, Niyati Malhotra, Sharif Kazemi, Victor Orozco-Olvera, Ana Mar\'ia Mu\~noz Boudet, Lakshmi Subramanian, Samuel P. Fraiberger, Sharath Chandra Guntuku, Valentin Hofmann