The paper investigates the nature of agreement among repeated samples of large language models (LLMs), showing that strong agreement can arise even for incorrect answers. It introduces a pluralistic agreement index, Gamma, which is decomposed into a mechanical component driven solely by per‑case answer preferences and a residual component that captures preference‑unexplained agreement. Experiments on GPT‑4.1 and several open‑weight models demonstrate that mechanical agreement dominates in many settings, while the residual varies with benchmark type and sampling protocol.
By Lizhuo Zhang, Mengmeng Tang, Chenfeng Long, Xiaoyong Tang, Xiang Luo
arXiv:2606. 03003v1 Announce Type: cross Abstract: A latent world model built from an equivariant encoder $E$ and an equivariant predictor $f$ inherits a provable symmetry of its training loss: when the world's dynamics genuinely carries a group $G$ acting on latents by an orthogonal representation $\rho(g)$, the one-step prediction relMSE is exactly invariant across the whole group, so fitting the dynamics on a restricted slice of orientations mathematically determines it on the entire orbit (j\v{u} y\=i f\v{a}n s\=an).
By Hongbo Wang (Stony Brook University)
The paper introduces a pluralistic agreement index, Gamma, to quantify how often wrong runs of large language models (LLMs) agree with the majority consensus. By decomposing Gamma into a mechanical component and a preference‑unexplained residual, the authors show that on GPT‑4.1 the mechanical part explains most of the agreement on multiple‑choice benchmarks but only about half on open‑domain tasks, revealing a residual bias that can cause self‑consistency to backfire on hard questions. The study provides a quantitative framework for understanding when majority voting over LLM samples improves or harms accuracy, without proposing new voting methods.
By Lizhuo Zhang, Mengmeng Tang, Chenfeng Long, Xiaoyong Tang, Xiang Luo
The paper evaluates additive activation steering in chat and agent contexts, showing that the commonly used gain ratio (Δ_agent/Δ_chat) fails to reliably indicate potency and efficacy across multiple models and dose-response cells. By replacing the gain with a location metric, dEC50 (difference in EC50 between agent and chat), the authors demonstrate a more robust, two‑sided measure that consistently captures cross‑context shifts. The study also reports several refuted and unanswered claims, emphasizing that a single operating point cannot distinguish between displacement and gain effects.
By Lucas Pinto
arXiv:2607. 17047v1 Announce Type: cross Abstract: LLM constraint reasoners are often evaluated near the random-SAT phase transition, confounding density and solver hardness.
By Lucky Verma
The paper introduces a site-asymmetry audit for activation-space interventions, decomposing order-dependent activation statistics into a canonical additive response from single interventions and an antisymmetrized second difference that removes first-order and self-curvature effects. Across six language-model families, the single-intervention baseline accounts for most of the bracket norm, while the corrected residual often clears a generic-interaction null. The method also transfers to non-language models, demonstrating its portability.
By Anqi Peter Li