arXiv:2607. 17531v1 Announce Type: cross Abstract: Test-time collaboration, including self-consistency, best-of-N selection, critic models, and verifier pipelines, is often credited with broadly improving LLM reasoning, yet its gains are uneven and sometimes negative.
By Jie Hu
arXiv:2609.27822v1 Announce Type: cross
Abstract: A common multi-agent design asks agents to report confidence and lets the highest-scoring agent speak next, implicitly using one scalar both to route...
By Jingyan Jiang, Huihuo Zheng, Rajeev Thakur, Chih-Hsuan Yang
arXiv:2608. 08265v1 Announce Type: new Abstract: Oracle routing measures how much a pool of language models could gain from per-query selection, but the diagnostic has two flaws: testing against a best fixed model selected on the same examples invalidates paired inference, and a full-information oracle sees outcomes no deployable router observes.
By Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb
COMED (Controlled Model Escalation for Multi-LLM Deliberation) is a post-anchor controller that selectively engages cross-model collaboration in multi-LLM inference. It uses anchor self‑consistency, router margin, and a lightweight peer probe to accept confident answers, verify ambiguous cases, and only escalates when collaboration is likely beneficial. Experiments on medical, scientific, and general reasoning benchmarks show that COMED improves performance across 16 open‑weight settings, achieving up to +10.7 percentage points on MedQA and outperforming dense collaboration while invoking fewer models and decoded tokens.
By Norah Alballa, Wenxuan Zhang, Salma Kharrat, Fares Fourati, Zafar Ayyub Qazi, Mohamed Elhoseiny, Marco Canini
MeshHeal is a fully decentralized self‑healing framework for decentralized LLM‑based multi‑agent systems that addresses gray failures—situations where an agent remains responsive but its task‑solving quality degrades. It operates on two timescales: a fast adaptive hierarchy that escalates uncertain or low‑scoring outputs to committee review and correction, and a slow peer‑relative detector that aggregates scores to distinguish persistent degradation from normal variation, triggering mandatory review and eventual exclusion of degraded agents while allowing recovered agents to rejoin. MeshHeal’s evaluation, using Model‑Backed MAS Evaluation, shows it achieves higher degraded‑phase accuracy (0.839) on BBH, MATH, and MMLU‑Pro with fewer tokens per task compared to the baseline Symphony.
By Keru Chen, Sen Lin, Yingbin Liang, Nathaniel D. Bastian, Shaofeng Zou
arXiv:2608. 04804v1 Announce Type: cross Abstract: Frontier language models can resolve repository-level software issues, but each attempt is expensive, and existing routers select a model from the issue text alone.
By Ishaan Bhola, Adithyan Krishnan, Mukunda NS