The study investigates how varying the capability of a reviewer model in a large‑language‑model (LLM) execute‑review‑revise pipeline affects rejection decisions on 100 olympiad mathematics problems. A mid‑tier reviewer improves final accuracy by 12 percentage points (from 52 % to 64 %) without damaging answers, while a self‑reviewer detects errors best (85 % recall) but rejects too often and harms correct solutions. Below a certain capability threshold the reviewer becomes inert, changing none of the answers and doubling token cost.
By Faizan Tanveer
MeshHeal is a fully decentralized self‑healing framework for decentralized LLM‑based multi‑agent systems that addresses gray failures—situations where an agent remains responsive but its task‑solving quality degrades. It operates on two timescales: a fast adaptive hierarchy that escalates uncertain or low‑scoring outputs to committee review and correction, and a slow peer‑relative detector that aggregates scores to distinguish persistent degradation from normal variation, triggering mandatory review and eventual exclusion of degraded agents while allowing recovered agents to rejoin. MeshHeal’s evaluation, using Model‑Backed MAS Evaluation, shows it achieves higher degraded‑phase accuracy (0.839) on BBH, MATH, and MMLU‑Pro with fewer tokens per task compared to the baseline Symphony.
By Keru Chen, Sen Lin, Yingbin Liang, Nathaniel D. Bastian, Shaofeng Zou
arXiv:2610.01023v1 Announce Type: cross
Abstract: Coding agents can return plausible patches that omit required behavior. These failures are hard to review because long traces and confident summaries...
By Junyu Guo, Shangding Gu, Ming Jin, Javad Lavaei
arXiv:2608.21374v1 Announce Type: new
Abstract: Literature reviews are essential to scientific progress, but rigorously evaluating automatically generated reviews remains difficult because many aspec...
By Ruotong Zhao, Zhiyu Chen, Xurui Liu, Haidong Xue, Dong Liang, Jigao Fu, Wu YanBiao, Yuanyi Zhen, Fengli Xu, Yong Li
arXiv:2607. 01223v1 Announce Type: new Abstract: When should an AI system's answer be trusted?
By Ben Slivinski, Michael Saldivar
arXiv:2608. 14927v1 Announce Type: new Abstract: Multi-agent large language model (LLM) systems can improve reasoning by spending more computation, but deployment requires deciding when extra collaboration is worth its cost.
By Chih-Hsuan Yang, Jingyan Jiang, Cheng-Hau Yang, Vikram Vasudevan, Huihuo Zheng, Venkatram Vishwanath, Rajeev Thakur
The paper introduces Independent–Communicate–Revise (ICR), a framework that isolates communication effects in large language model multi‑agent systems by fixing initial reasoning and measuring how messages influence answer revision. ICR evaluates correction, preservation, and selectivity across four reasoning benchmarks, revealing that similar overall accuracy can mask divergent revision behaviors. The study shows that richer messages can both improve and harm outcomes, and that receiver policies can shift preservation and correction dynamics differently across tasks.
By Shixuan Li, Wei Yang, Peiyu Zhang, Anzhe Cheng, Heng Ping, Paul Bogdan
arXiv:2604. 16706v2 Announce Type: replace Abstract: Automated evaluation of tool-using large language model (LLM) agents is widely assumed to be reliable, yet this assumption is rarely validated against human annotation.
By Bhaskar Gurram
Adversarial Review (AR) is a minimal cooperative code‑review protocol that employs a main coding agent, a reviewer, and a critic. The reviewer evaluates code while the critic audits the review through structured disagreement before the main agent edits. On multiple benchmarks (LiveCodeBench, SWE‑PRBench, SWE‑bench Verified), AR achieves higher pass rates or F1 scores than larger multi‑agent baselines, demonstrating that effective code review can be achieved with only three agents and minimal, evidence‑grounded disagreement.
By Eric S. Qiu, Joyce Gill
Ideation Arena is a battle-style platform that evaluates research ideas generated by large language models (LLMs) and research agents through pairwise human assessment. The system builds shared literature contexts, collects over 6,000 double-blind comparisons from 105 computer science researchers, and constructs an Elo rating leaderboard to rank proposal-stage expert preferences. It also introduces Ideation Arena Eval, a benchmark to test whether automated evaluators align with human preferences, finding that current LLM judges achieve at best 72.56% Soft Accuracy on overall quality.
By Zhiyu Chen, Keyu Zhao, Jigao Fu, Dong Liang, Yanbiao Wu, Jiaoyang Li, Haidong Xue, Xinhua Zeng, Yuanyi Zhen, Fengli Xu, Yong Li
arXiv:2607. 08065v1 Announce Type: new Abstract: LLM-as-judge (Zheng et al.
By Kaihua Ding
arXiv:2610.00885v1 Announce Type: cross
Abstract: Coding agents increasingly automate Lean proof development, but successful compilation alone does not establish that a candidate proves the intended...
By Naing Oo Lwin