arXiv AI

Decomposing Wrong-Consensus Agreement in LLM Self-Consistency

The paper investigates the nature of agreement among repeated samples of large language models (LLMs), showing that strong agreement can arise even for incorrect answers. It introduces a pluralistic agreement index, Gamma, which is decomposed into a mechanical component driven solely by per‑case answer preferences and a residual component that captures preference‑unexplained agreement. Experiments on GPT‑4.1 and several open‑weight models demonstrate that mechanical agreement dominates in many settings, while the residual varies with benchmark type and sampling protocol.

arXiv AI
Aug 20

Decomposing Wrong-Consensus Agreement in LLM Self-Consistency: A GPT-4.1 Case Study

The paper introduces a pluralistic agreement index, Gamma, to quantify how often wrong runs of large language models (LLMs) agree with the majority consensus. By decomposing Gamma into a mechanical component and a preference‑unexplained residual, the authors show that on GPT‑4.1 the mechanical part explains most of the agreement on multiple‑choice benchmarks but only about half on open‑domain tasks, revealing a residual bias that can cause self‑consistency to backfire on hard questions. The study provides a quantitative framework for understanding when majority voting over LLM samples improves or harms accuracy, without proposing new voting methods.

By Lizhuo Zhang, Mengmeng Tang, Chenfeng Long, Xiaoyong Tang, Xiang Luo
arXiv Machine Learning
Jun 2

Measuring the Symmetry--Data Exchange Rate

arXiv:2606. 01090v1 Announce Type: cross Abstract: Equivariance theory predicts that an architectural symmetry prior reduces sample complexity by a factor of |G|; this is widely cited but rarely measured as a scaling law with controls that separate the prior from its confounds.

By Ahmed M. Adly