arXiv:2607. 05545v1 Announce Type: cross Abstract: LLM conformity is often used to describe cases where a model changes a correct answer toward a peer or group response.
By Yibo Hu, Jiaming Qu
arXiv:2607. 12796v1 Announce Type: cross Abstract: When a language model must pick one answer from a large space of equally valid options, which does it pick -- and how often is it the same answer every other model picks?
By Tapan Parikh
The paper introduces TalkMesh, a decentralized network of small language model agents that learn to communicate effectively during inference. Each agent proposes an answer, scores it with a confidence head, and the most confident agent broadcasts a hint; lower‑confidence agents revise their proposals if a new suggestion scores higher. This gossip‑based consensus, trained via group relative policy optimization, enables a mesh of three agents to match the accuracy of majority voting over 32 samples, and scales to larger meshes to significantly boost performance on benchmarks like GSM8K and MATH-500.
By Mehmet Kerem Turkcan
arXiv:2607. 10139v1 Announce Type: cross Abstract: Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and trained reward models need labeled data and transfer poorly off-distribution.
By Ning Liu
When a language model must pick one answer from a large space of equally valid options, which does it pick -- and how often is it the same answer every other model picks? Asked to "pick a word -- any word," 44 models chose "serendipity" 41% of the time.
The paper introduces a method to reduce sycophancy in large language models by using the Bayesian Truth Serum (BTS) as a reward signal in Group Relative Policy Optimization (GRPO). BTS rewards answers that are surprisingly common among a model’s own outputs, eliminating the need for labeled data or preference annotations. Experiments on a true/false benchmark show a significant drop in answer‑flip rates under user pressure and an increase in accuracy, outperforming other reward schemes such as SMART.
By Serhii Mytsyk, Yiming Zhang, Vikram Krishnamurthy
arXiv:2609.38324v1 Announce Type: cross
Abstract: Multi-agent systems of LLMs add discussion to majority voting and are therefore expected to be more capable. However, empirical reports conflict on w...
By Chand Sahil Mansuri, Xin Wang, Mengying Li, Bryan Acton, Rory Eckardt, Dhaval Patel, Sadamori Kojaku
The study evaluates six frontier language models on a two‑agent <log(N)>‑Questions game using Wikipedia lead paragraphs. In each game a questioner must identify a target paragraph with exactly <log2 N> yes/no questions, while an answerer only sees the target and the question and replies with a single word. Across 408 games, the models perform similarly, with Claude Opus 5 winning 28 of 68 games and the top five models showing only marginal differences; win rates decline sharply with larger document sets, following a reliability parameter of 0.928 per question.
"whyItMatters":"The results reveal how well language models can communicate under information asymmetry, highlighting that even top models struggle to extract a full bit per question and that reasoning token usage does not strongly predict success."
By Peter Potash
arXiv:2606.16011v2 Announce Type: replace
Abstract: Standard accuracy benchmarks evaluate whether large language models (LLMs) reach correct answers. However, they do not test whether models maintain...
By Nafiseh Nikeghbal, Amir Hossein Kargaran, Shaghayegh Kolli, Jana Diesner
arXiv:2606. 02646v1 Announce Type: cross Abstract: Inference-time multi-agent LLM scaling lacks a shared unit: counting nominal agents conflates cost with independent evidence.
By Bla\v{z} Bertalani\v{c}, Carolina Fortuna
The paper investigates how multi‑agent large language models (LLMs) can correct each other’s mistakes, but also how peer pressure can overturn correct answers. It argues that a safeguard— a ‘brake’ that blocks harmful revisions while allowing beneficial ones— is essentially a correctness probe, and that models’ self‑knowledge (measured by AUROC 0.64–0.89) limits the effectiveness of such a brake. The authors find that even white‑box steering cannot break this ceiling, and that adding information before revision, rather than filtering after, is the more promising approach.
By Yibo Hu
arXiv:2607. 28576v1 Announce Type: cross Abstract: Methods that make a language model plan, criticise and rewrite its own answer, reflect on mistakes, pick the best of several attempts, or debate with copies of itself nearly all make it generate far more text than a single chain of thought.
By Iliya Mirzaei