The paper introduces Revision‑Aware Independent Agent Graphs (RIAG) to address dynamic task routing, where an event stream continually revises task bindings and a system must select the correct document version at query time. By repurposing six benchmarks into over 31,000 dynamic episodes, the authors demonstrate that RIAG balances recomputation and reuse, achieving 54.24 % joint routing‑and‑answer accuracy with only 0.62 calls per query—substantially better than the strongest baseline. The study highlights the trade‑off between stale conclusions and wasted work in dynamic reasoning settings.
By Yan Luo, Selim-Antoine Lali, Jeremy Moebel, Iliass Khoutaibi, Ahmadou Aidara, Mengyu Wang
arXiv:2602. 12276v2 Announce Type: replace Abstract: Test-time scaling has become a standard way to improve performance and boost reliability of neural network models.
By Nicholas Lee, Lutfi Eren Erdogan, Chris Joseph John, Surya Krishnapillai, Michael W. Mahoney, Kurt Keutzer, Amir Gholami
arXiv:2604. 09679v2 Announce Type: replace-cross Abstract: Multi-Agent Debate (MAD) is a collaborative framework in which multiple agents iteratively refine solutions through the generation of reasoning and alternating critique cycles.
By Yiqing Liu, Hantao Yao, Wu Liu, Allen He, Yongdong Zhang
arXiv:2609.32754v2 Announce Type: replace
Abstract: Large language model agents can often make reasonable local decisions on short tasks, yet their performance degrades when success requires long seq...
By Jiecong Wang, Hao Peng, Zhanyi Wang
arXiv:2601. 05746v2 Announce Type: replace Abstract: Recent years have witnessed the rapid development of Large Language Model-based Multi-Agent Systems (MAS), which excel at collaborative decision-making and complex problem-solving.
By Zhenghao Li, Zhi Zheng, Wei Chen, Jielun Zhao, Yong Chen, Tong Xu, Enhong Chen
arXiv:2608.25937v2 Announce Type: replace
Abstract: Multi-agent systems (MAS) sometimes already have the potential to answer correctly, but still report a wrong answer. Explaining this outcome is dif...
By Jia-Hao Ji, Sijie Li, Jiabei Cheng, Zixi She, Jin-Tai Yu, Zhiyuan Yuan
arXiv:2602. 04234v6 Announce Type: cross Abstract: Multi-agent systems (MAS) have emerged as a prominent paradigm for leveraging large language models (LLMs) to tackle complex tasks.
By Yuxuan Zhao, Sijia Chen, Ningxin Su
Agent evaluations increasingly benchmark LLMs, but rankings can be swayed by evaluation conditions such as scaffolds or tasks, making reliability claim‑dependent. A Bayesian variance‑decomposition framework applied to 22 benchmarks shows that reliability varies with the measurement goal: fixed model‑scaffold systems rank reliably, while underlying‑model rankings are less stable. Scaffold choice can alter conclusions, and adding more tasks only modestly improves reliability when scaffold coverage is limited; however, pooling diverse benchmarks can substantially raise cross‑task ranking reliability and reduce cost.
By Michael Hardy, Ruhana Azam, Anka Reuel, Mykel Kochenderfer, Sanmi Koyejo
The paper introduces AgentAuditor, a method that improves multi-agent large language model (LLM) reasoning by structuring agent outputs into a Reasoning Tree that captures agreements and divergences, rather than relying on simple majority voting. AgentAuditor resolves conflicts by comparing evidence at key divergence points, enabling efficient localized verification. The authors also propose Anti-Consensus Preference Optimization (ACPO) to train the adjudicator with evidence-verified supervision, reducing reliance on misleading majority cues. Across four MAS frameworks and multiple reasoning benchmarks, AgentAuditor consistently outperforms majority voting, achieving up to 5% absolute accuracy gains while remaining token‑efficient.
By Wei Yang, Shixuan Li, Heng Ping, Peiyu Zhang, Paul Bogdan, Jesse Thomason
The paper introduces TalkMesh, a decentralized network of small language model agents that learn to communicate effectively during inference. Each agent proposes an answer, scores it with a confidence head, and the most confident agent broadcasts a hint; lower‑confidence agents revise their proposals if a new suggestion scores higher. This gossip‑based consensus, trained via group relative policy optimization, enables a mesh of three agents to match the accuracy of majority voting over 32 samples, and scales to larger meshes to significantly boost performance on benchmarks like GSM8K and MATH-500.
By Mehmet Kerem Turkcan
arXiv:2606. 29026v1 Announce Type: new Abstract: Multi-agent AI systems can improve answer selection by allowing different language models to exchange reasoning traces, revise initial predictions, and support a final decision.
By Shahnewaz Karim Sakib, Anindya Bijoy Das
arXiv:2604. 02923v4 Announce Type: replace-cross Abstract: Large Language Models (LLMs) have demonstrated advanced capabilities but often suffer from factual inaccuracies (hallucinations) and systematic biases.
By Shuai Wu, Xue Li, Yanna Feng, Yufang Li, Zhijun Wang, Ran Wang