arXiv:2607. 17531v1 Announce Type: cross Abstract: Test-time collaboration, including self-consistency, best-of-N selection, critic models, and verifier pipelines, is often credited with broadly improving LLM reasoning, yet its gains are uneven and sometimes negative.
By Jie Hu
arXiv:2609.27822v1 Announce Type: cross
Abstract: A common multi-agent design asks agents to report confidence and lets the highest-scoring agent speak next, implicitly using one scalar both to route...
By Jingyan Jiang, Huihuo Zheng, Rajeev Thakur, Chih-Hsuan Yang
arXiv:2608. 08265v1 Announce Type: new Abstract: Oracle routing measures how much a pool of language models could gain from per-query selection, but the diagnostic has two flaws: testing against a best fixed model selected on the same examples invalidates paired inference, and a full-information oracle sees outcomes no deployable router observes.
By Ibne Farabi Shihab, Abu Sa-Adat Mohamed Moon-Im Al Ahsan, Md Najmus Swaqeeb
COMED (Controlled Model Escalation for Multi-LLM Deliberation) is a post-anchor controller that selectively engages cross-model collaboration in multi-LLM inference. It uses anchor self‑consistency, router margin, and a lightweight peer probe to accept confident answers, verify ambiguous cases, and only escalates when collaboration is likely beneficial. Experiments on medical, scientific, and general reasoning benchmarks show that COMED improves performance across 16 open‑weight settings, achieving up to +10.7 percentage points on MedQA and outperforming dense collaboration while invoking fewer models and decoded tokens.
By Norah Alballa, Wenxuan Zhang, Salma Kharrat, Fares Fourati, Zafar Ayyub Qazi, Mohamed Elhoseiny, Marco Canini
MeshHeal is a fully decentralized self‑healing framework for decentralized LLM‑based multi‑agent systems that addresses gray failures—situations where an agent remains responsive but its task‑solving quality degrades. It operates on two timescales: a fast adaptive hierarchy that escalates uncertain or low‑scoring outputs to committee review and correction, and a slow peer‑relative detector that aggregates scores to distinguish persistent degradation from normal variation, triggering mandatory review and eventual exclusion of degraded agents while allowing recovered agents to rejoin. MeshHeal’s evaluation, using Model‑Backed MAS Evaluation, shows it achieves higher degraded‑phase accuracy (0.839) on BBH, MATH, and MMLU‑Pro with fewer tokens per task compared to the baseline Symphony.
By Keru Chen, Sen Lin, Yingbin Liang, Nathaniel D. Bastian, Shaofeng Zou
arXiv:2608. 04804v1 Announce Type: cross Abstract: Frontier language models can resolve repository-level software issues, but each attempt is expensive, and existing routers select a model from the issue text alone.
By Ishaan Bhola, Adithyan Krishnan, Mukunda NS
arXiv:2606. 27288v1 Announce Type: new Abstract: Multi-model LLM systems such as routing, voting, cascades, fusion, and mixture-of-agents are used to beat single-model accuracy.
By Josef Chen
arXiv:2605.18859v3 Announce Type: replace-cross
Abstract: LLM routing matters most in long-horizon applications such as coding agents, deep research systems, and computer-use agents, where a single u...
By Pei Yang, Wanyi Chen, Tongyun Yang, Pengbin Feng, Jiarong Xing, Wentao Guo, Yuhang Yao, Yuhang Han, Hanchen Li, Xu Wang, Zeyu Wang, Jie Xiao, Anjie Yang, Liang Tian, Lynn Ai, Eric Yang, Tianyu Shi
arXiv:2608.23023v2 Announce Type: new
Abstract: An LLM router picks which model should answer each query. The appeal is that models fail on different questions. Whatever single model is best overall...
By Janghoon Lee (Redrob)
The paper investigates why large‑language‑model (LLM) routers—systems that select the best model for each query—often fail to outperform a single best model. By evaluating 14 models on 294 questions across seven task types and three languages, the authors find that a simple static mapping of task type to model improves 21 of the 29 questions that routing could potentially solve, and that learned routers do not significantly exceed this performance. The study highlights that most routing gains stem from task‑type specialization rather than complex learned decision rules.
By Janghoon Lee
arXiv:2607. 15388v1 Announce Type: new Abstract: Many math- and science-oriented agent systems use hierarchical designs with specialized reviewer roles, assuming that a dedicated review stage should help turn wrong candidates into correct ones.
By Chih-Hsuan Yang, Jingyan Jiang, Vikram Vasudevan, Cheng-Hau Yang, Huihuo Zheng, Le Chen, Eliu A. Huerta, Venkatram Vishwanath, Ian T. Foster, Rajeev Thakur
The paper introduces a two‑level readout for mixture‑of‑experts reasoning models. First, it compresses the model’s internal reasoning states into a 64‑dimensional semantic frame (J64) that reveals process dynamics beyond the emitted trace. Second, it reconstructs this frame from native expert‑routing statistics (R64), achieving high correlation and preserving most predictive gains while enabling low‑overhead, test‑time decision making.
By Kang Chen, Sihan Zhao, Yixin Cao, Yugang Jiang