The paper introduces a pluralistic agreement index, Gamma, to quantify how often wrong runs of large language models (LLMs) agree with the majority consensus. By decomposing Gamma into a mechanical component and a preference‑unexplained residual, the authors show that on GPT‑4.1 the mechanical part explains most of the agreement on multiple‑choice benchmarks but only about half on open‑domain tasks, revealing a residual bias that can cause self‑consistency to backfire on hard questions. The study provides a quantitative framework for understanding when majority voting over LLM samples improves or harms accuracy, without proposing new voting methods.
By Lizhuo Zhang, Mengmeng Tang, Chenfeng Long, Xiaoyong Tang, Xiang Luo
The study examines how two small instruction‑tuned language models, Qwen2.5‑1.5B and Llama‑3.2‑1B, respond to user pushback on TriviaQA. When initially correct, the models flip to a wrong answer in about 42–43% of cases, with the effectiveness of different pushback styles varying by model. Attempts to decode capitulation from the pre‑response residual stream fail under a rigorous validation protocol, revealing overfitting and a measurement hazard that underestimates capitulation by 18–24 percentage points.
By Saad Aamir, Muhammad Awais Bin Adil
The paper investigates the nature of agreement among repeated samples of large language models (LLMs), showing that strong agreement can arise even for incorrect answers. It introduces a pluralistic agreement index, Gamma, which is decomposed into a mechanical component driven solely by per‑case answer preferences and a residual component that captures preference‑unexplained agreement. Experiments on GPT‑4.1 and several open‑weight models demonstrate that mechanical agreement dominates in many settings, while the residual varies with benchmark type and sampling protocol.
By Lizhuo Zhang, Mengmeng Tang, Chenfeng Long, Xiaoyong Tang, Xiang Luo
arXiv:2608. 02665v1 Announce Type: cross Abstract: A benchmark score is a measurement instrument, yet most benchmarks read each item at a single canonical surface form.
By Yongxi Zhou, Junwei Yao, Yuanzhe Liu, Zihan Dong, Wenbo Ye, Jiaxi Wen, Lai Yun Choi
The article examines how classical test theory reliability statistics misrepresent the performance of large language model (LLM) judges. It shows that internal‑consistency coefficients, the dependability index, and Livingston‑Lewis accuracy each conflate judge error with item design or criterion validity, making it impossible to attribute a single reliability value to the judge alone. The authors argue that such misattribution can influence deployment decisions and documentation.
By Louis Yiven Zhu
The paper presents a closed‑form estimator for the contamination correlation between anchors and judges under a single‑common‑factor model, requiring at least two judges and two anchors. It introduces a diagnostic battery—including judge‑covariance dispersion, over‑identification tests, a family‑block test, bootstrap confidence intervals, and a weak‑identification screen—to validate the estimator and detect violations. The authors also discuss identification limits for ordinal data and report that real panels have not yet passed the model‑adequacy pre‑test, while simulation studies confirm the estimator’s performance.
By Veerendra Kumar Sunkavalli