arXiv:2606. 01637v1 Announce Type: cross Abstract: Large language models are increasingly used in multi-agent systems, where they see and respond to other agents' answers.
By Jiaming Qu, Lucheng fu, Yibo Hu
The paper investigates how multi‑agent large language models (LLMs) can correct each other’s mistakes, but also how peer pressure can overturn correct answers. It argues that a safeguard— a ‘brake’ that blocks harmful revisions while allowing beneficial ones— is essentially a correctness probe, and that models’ self‑knowledge (measured by AUROC 0.64–0.89) limits the effectiveness of such a brake. The authors find that even white‑box steering cannot break this ceiling, and that adding information before revision, rather than filtering after, is the more promising approach.
By Yibo Hu
arXiv:2607. 05545v1 Announce Type: cross Abstract: LLM conformity is often used to describe cases where a model changes a correct answer toward a peer or group response.
By Yibo Hu, Jiaming Qu
arXiv:2608. 11247v1 Announce Type: new Abstract: Recent advances in language models have enabled collaborative settings in which multiple models leverage one another's capabilities, iteratively improving, transforming, and extending each other's outputs.
By Zafar Hussain, Kristoffer Nielbo
arXiv:2608.25937v2 Announce Type: replace
Abstract: Multi-agent systems (MAS) sometimes already have the potential to answer correctly, but still report a wrong answer. Explaining this outcome is dif...
By Jia-Hao Ji, Sijie Li, Jiabei Cheng, Zixi She, Jin-Tai Yu, Zhiyuan Yuan
This paper introduces a Byzantine‑robust federated retrieval‑augmented generation (RAG) framework that uses aligned calibration and fixed‑membership conformal prediction to ensure that the answer set contains the correct answer with a chosen probability, even when some nodes are compromised. By having all nodes score the same calibration questions and retaining only candidates that could be kept by a plausible group of honest nodes, the method guarantees correctness in finite samples and produces smaller answer sets than simpler approaches. Experiments on medical exam question‑answering tasks with language‑model nodes demonstrate that the method meets the target coverage whenever the number of misbehaving nodes does not exceed the declared bound, while plain averaging often fails.
By Prasanjit Dubey, Aritra Guha, Xiaoming Huo
arXiv:2607. 24780v1 Announce Type: new Abstract: Evaluating frontier LLMs is challenging: static benchmarks suffer from contamination and saturation -- leaving users unable to distinguish top models and developers blind to specific failure modes -- while human preference is subjective.
By Xingyu Chen, Rui Wang, Zhaopeng Tu, Liefeng Bo
arXiv:2609.37493v1 Announce Type: cross
Abstract: Serving an answer from a large language model requires deciding when to abstain, yet a verifier's ranking accuracy alone does not determine the error...
By Dongyub Jude Lee, Jungseob Lee, Chanjun Park, Hyeonseok Moon, Heuiseok Lim
arXiv:2606. 29054v1 Announce Type: new Abstract: Large language models (LLMs) deployed for structured generation (NER, JSON extraction, QA, and classification) lack formal reliability guarantees, and standard heuristic abstention policies miss user-specified risk targets by 7.
By Varun Kotte
arXiv:2607. 08456v1 Announce Type: cross Abstract: A model should refuse two different things: answers it would get wrong, and questions it should not answer at all, such as unanswerable ones or ones resting on a false premise.
By Benedikt J. Wagner
arXiv:2608.23638v1 Announce Type: cross
Abstract: Two analysts who calibrate the same predictive model on independent samples will deploy different prediction sets every time, because the calibration...
By Marios Papamichalis, Regina Ruane, Theofanis Papamichalis
arXiv:2606. 29654v1 Announce Type: new Abstract: Multi-agent deliberation among LLMs can improve reasoning, but deployment requires deciding when the current answer is reliable enough to act on and when it should be escalated to human review.
By Mengdie Flora Wang, Haochen Xie, Guanghui Wang, Devin Zhang, Jae Oh Woo