Hugging Face Trending Papers

Preference Is Not Intervention: The Structure and Stability Boundaries of Reader-Specific Evidence Utility

Read the original on Hugging Face Trending Papers →

The paper investigates whether reader-specific differences in retrieval‑augmented generation (RAG) reflect reusable structure or merely input‑local interactions. By fixing query, evidence, task, scoring, and intervention, the authors find that nine readers disagree on the effect sign in 33% of cases, with reader×query interactions explaining 29.8% of utility variance. They further decompose heterogeneity into evidence activity, ordinal preference, and conditional signed direction, discovering that ordinal reader geometry is stable across multiple settings while signed geometry is task‑bounded, yet stable ordinal similarity does not predict cross‑reader intervention transfer.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Aug 26

How much of a measured AI preference is the model, and how much is the instrument?

The paper investigates how much of an AI model’s expressed preferences are due to the model itself versus the instrument (prompt) used to elicit those preferences. By fixing the set of outcomes and models while varying five different prompting instruments across 15 welfare-related outcomes, the study finds that the ranking of outcomes is only moderately generalizable (coefficient 0.348) and that a single instrument’s preference provides little insight into another instrument’s results. The analysis shows that even after removing any single instrument, model, or a subset of outcomes, the overall preference estimate remains robust, yet the variability across instruments remains significant.

By Jason Hung
arXiv AI
Aug 20

Decomposing Wrong-Consensus Agreement in LLM Self-Consistency: A GPT-4.1 Case Study

The paper introduces a pluralistic agreement index, Gamma, to quantify how often wrong runs of large language models (LLMs) agree with the majority consensus. By decomposing Gamma into a mechanical component and a preference‑unexplained residual, the authors show that on GPT‑4.1 the mechanical part explains most of the agreement on multiple‑choice benchmarks but only about half on open‑domain tasks, revealing a residual bias that can cause self‑consistency to backfire on hard questions. The study provides a quantitative framework for understanding when majority voting over LLM samples improves or harms accuracy, without proposing new voting methods.

By Lizhuo Zhang, Mengmeng Tang, Chenfeng Long, Xiaoyong Tang, Xiang Luo
arXiv Machine Learning
Sep 22

Resist, Update, Reject: Preference Optimization Installs a Prior-Dependent Reliability Switch

The paper demonstrates that a preference‑optimization objective can learn to distinguish reliable from unreliable sources by installing a prior‑dependent reliability switch. By training on data where a source’s stated reliability is paired with its answer, the model learns to flip its response only when the stated reliability exceeds a threshold that grows with the model’s prior. Experiments on Qwen2.5‑7B‑Instruct and Llama‑3.1‑8B show that this switch generalizes to unseen reliability values and follows stated reliability over role prestige, whereas supervised imitation fails to learn it.

By Sen Yang, Yuen-Hei Yeung
arXiv AI
Sep 24

Ask Which, Not How Good: Sizing Benchmarks Scored by an LLM

The study analyzes 373,019 judgments from LLM‑scored benchmarks, decomposing variance into system, item, judge, and interaction components via generalizability theory. It finds that with a single judge, generalizability converges to a ceiling determined by the system‑by‑judge variance, which is substantially lower in pairwise preference settings, allowing one judge to suffice. The research also reveals significant biases in presentation order and highlights that many published win‑rate claims fall below the measured floor of the benchmarks.

By Atul Anand