arXiv:2607. 05545v1 Announce Type: cross Abstract: LLM conformity is often used to describe cases where a model changes a correct answer toward a peer or group response.
By Yibo Hu, Jiaming Qu
The paper investigates how multi‑agent large language models (LLMs) can correct each other’s mistakes, but also how peer pressure can overturn correct answers. It argues that a safeguard— a ‘brake’ that blocks harmful revisions while allowing beneficial ones— is essentially a correctness probe, and that models’ self‑knowledge (measured by AUROC 0.64–0.89) limits the effectiveness of such a brake. The authors find that even white‑box steering cannot break this ceiling, and that adding information before revision, rather than filtering after, is the more promising approach.
By Yibo Hu
The paper introduces Independent–Communicate–Revise (ICR), a framework that isolates communication effects in large language model multi‑agent systems by fixing initial reasoning and measuring how messages influence answer revision. ICR evaluates correction, preservation, and selectivity across four reasoning benchmarks, revealing that similar overall accuracy can mask divergent revision behaviors. The study shows that richer messages can both improve and harm outcomes, and that receiver policies can shift preservation and correction dynamics differently across tasks.
By Shixuan Li, Wei Yang, Peiyu Zhang, Anzhe Cheng, Heng Ping, Paul Bogdan
arXiv:2609.36855v1 Announce Type: new
Abstract: Multi-agent LLM systems rely on message passing among specialized agents to accomplish complex tasks. However, an upstream agent may provide useful inf...
By Yaxin Gong, Gangyi Zhang, Chongming Gao, Leyang Shen, Chenxiao Fan, Jiakai Wang, Dong Wang, Yang Liu, Wenjie Wang, Xiangnan He
arXiv:2607. 24780v1 Announce Type: new Abstract: Evaluating frontier LLMs is challenging: static benchmarks suffer from contamination and saturation -- leaving users unable to distinguish top models and developers blind to specific failure modes -- while human preference is subjective.
By Xingyu Chen, Rui Wang, Zhaopeng Tu, Liefeng Bo
arXiv:2609.38324v1 Announce Type: cross
Abstract: Multi-agent systems of LLMs add discussion to majority voting and are therefore expected to be more capable. However, empirical reports conflict on w...
By Chand Sahil Mansuri, Xin Wang, Mengying Li, Bryan Acton, Rory Eckardt, Dhaval Patel, Sadamori Kojaku
arXiv:2607. 28908v1 Announce Type: new Abstract: Reflection, the ability to revisit and revise prior reasoning, is central to how humans improve their answers.
By Yefan Tao, Gerald Friedland, Madhusudhanan Chandrasekaran, Luyang Kong
arXiv:2604. 02668v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often exhibit sycophancy: agreement with user stance even when it conflicts with the model's opinion.
By Vira Kasprova, Amruta Parulekar, Abdulrahman AlRabah, Krishna Agaram, Ritwik Garg, Sagar Jha, Nimet Beyza Bozdag, Dilek Hakkani-Tur
A new study shows that conformal certificates can become invalid when a large language model (LLM) is influenced by peers who unanimously provide a wrong answer, even though the question itself remains unchanged. This phenomenon, termed a score‑mechanism shift, reveals that a model’s calibration for single‑agent scoring does not hold in multi‑agent settings, leading to a drop in coverage from 90% to 74% under unanimous‑wrong peers. The shift also allows attackers to target low‑confidence items, nearly halving coverage for that subgroup while keeping overall averages deceptively high, and can cause systems to act confidently on incorrect answers.
whyItMatters":"The findings expose a critical vulnerability in conformal prediction for multi‑agent LLM systems, undermining their reliability and safety in real‑world applications."
By Yibo Hu, Hanyu Su
The paper introduces a novel framework for assessing second‑order bias in large language models (LLMs), defined as bias in how an LLM judges the acceptability of biased content. Using principles from entitlement epistemology, the authors design a reasoning task that asks LLMs to determine whether a biased text is acceptable for specific demographic groups, and propose two metrics to quantify biased judgments. Experiments on both open‑source and closed‑source models reveal that the task bypasses safety guardrails, uncovers systematic variations across target groups, and demonstrates that models still rely on demographic labels when evaluating bias.
By Ramaravind Kommiya Mothilal, Terry Jingchen Zhang, Raiyan Ahmed, Zhijing Jin, Shion Guha, Syed Ishtiaque Ahmed
The paper investigates intrinsic self‑correction, where a language model revises its own answer without new evidence. Across 29 open‑weight LLMs on BoolQ, GSM8K, and Corr2Cause, the study tracks how revisions change correctness, revealing that while some models improve significantly, others lose a notable fraction of correct answers. The authors compare three runtime strategies—keeping the initial answer, always accepting the revision, and selectively gating revisions—and find that the best approach depends on the model and task, suggesting that self‑correction should be treated as a revision policy rather than a uniformly beneficial second pass.
By Tianzhu Zhang
arXiv:2608.25937v2 Announce Type: replace
Abstract: Multi-agent systems (MAS) sometimes already have the potential to answer correctly, but still report a wrong answer. Explaining this outcome is dif...
By Jia-Hao Ji, Sijie Li, Jiabei Cheng, Zixi She, Jin-Tai Yu, Zhiyuan Yuan