arXiv AI

Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier

arXiv:2608. 11247v1 Announce Type: new Abstract: Recent advances in language models have enabled collaborative settings in which multiple models leverage one another's capabilities, iteratively improving, transforming, and extending each other's outputs.

arXiv Machine Learning
5d ago

Reinforcement Learning of Communication in a Mesh of Small Language Models

The paper introduces TalkMesh, a decentralized network of small language model agents that learn to communicate effectively during inference. Each agent proposes an answer, scores it with a confidence head, and the most confident agent broadcasts a hint; lower‑confidence agents revise their proposals if a new suggestion scores higher. This gossip‑based consensus, trained via group relative policy optimization, enables a mesh of three agents to match the accuracy of majority voting over 32 samples, and scales to larger meshes to significantly boost performance on benchmarks like GSM8K and MATH-500.

By Mehmet Kerem Turkcan
arXiv AI
Jul 14

LLMs as a Jury: Cross-Model Consensus Can Outperform Process Reward Models for LLM Reasoning

arXiv:2607. 10139v1 Announce Type: cross Abstract: Selecting the correct answer from a pool of candidate reasoning chains is the engine of test-time scaling, yet the standard selectors each carry a cost: self-consistency inherits the errors of the single model it resamples, and trained reward models need labeled data and transfer poorly off-distribution.

By Ning Liu
arXiv Machine Learning
Aug 27

Mitigating LLM sycophancy with RL-based fine-tuning: Bayesian Truth Serum approach

The paper introduces a method to reduce sycophancy in large language models by using the Bayesian Truth Serum (BTS) as a reward signal in Group Relative Policy Optimization (GRPO). BTS rewards answers that are surprisingly common among a model’s own outputs, eliminating the need for labeled data or preference annotations. Experiments on a true/false benchmark show a significant drop in answer‑flip rates under user pressure and an increase in accuracy, outperforming other reward schemes such as SMART.

By Serhii Mytsyk, Yiming Zhang, Vikram Krishnamurthy
arXiv Computation and Language
Sep 17

Playing log(N)-Questions over Wikipedia Abstracts: Communication Efficiency Between Paired Frontier Models

The study evaluates six frontier language models on a two‑agent <log(N)>‑Questions game using Wikipedia lead paragraphs. In each game a questioner must identify a target paragraph with exactly <log2 N> yes/no questions, while an answerer only sees the target and the question and replies with a single word. Across 408 games, the models perform similarly, with Claude Opus 5 winning 28 of 68 games and the top five models showing only marginal differences; win rates decline sharply with larger document sets, following a reliability parameter of 0.928 per question. "whyItMatters":"The results reveal how well language models can communicate under information asymmetry, highlighting that even top models struggle to extract a full bit per question and that reasoning token usage does not strongly predict success."

By Peter Potash
arXiv Machine Learning
Sep 17

One Axis, No Brake: Self-Knowledge Limits the Filtering of Harmful Peer Conformity in LLMs

The paper investigates how multi‑agent large language models (LLMs) can correct each other’s mistakes, but also how peer pressure can overturn correct answers. It argues that a safeguard— a ‘brake’ that blocks harmful revisions while allowing beneficial ones— is essentially a correctness probe, and that models’ self‑knowledge (measured by AUROC 0.64–0.89) limits the effectiveness of such a brake. The authors find that even white‑box steering cannot break this ceiling, and that adding information before revision, rather than filtering after, is the more promising approach.

By Yibo Hu