arXiv:2608. 08605v1 Announce Type: new Abstract: Multi-agent systems (MAS) built on Large Language Models (LLMs) are proliferating rapidly, but their heterogeneous execution traces provide no common basis for evaluation across methods.
By Guo Chen, Ziwen Li, Reed Li, Yu Lu, Haibo Shi, Bingbing Xu, Junjie Huang
arXiv:2601. 02854v2 Announce Type: replace Abstract: As an agent-level reasoning and coordination paradigm, Multi-Agent Debate (MAD) orchestrates multiple agents through structured debate to improve answer quality and support complex reasoning.
By Ao Li, Jinghui Zhang, Luyu Li, Yuxiang Duan, Lang Gao, Mingcai Chen, Weijun Qin, Shaopeng Li, Fengxian Ji, Ning Liu, Lizhen Cui, Xiuying Chen, Yuntao Du
Loom is a generative consensus framework designed for real‑world root‑cause analysis (RCA) that combines open‑form hypotheses from modular heuristics with a lightweight large language model (LLM) synthesis step. It projects hypotheses into a continuous embedding space and uses an iterative centroid‑based reweighting algorithm to resolve conflicts, producing a single consensus that is then synthesized by one LLM call. On the OpenRCA benchmark Loom matches state‑of‑the‑art autonomous agents on some datasets while achieving significantly higher efficiency—about 26× faster and 33× faster with an 8B‑parameter synthesizer.
whyItMatters":"Loom demonstrates how embedding‑space reweighting can bridge the gap between statistical rigor and expressive LLMs, enabling efficient, trustworthy RCA in industrial settings."
By Ron Begleiter, Katya Egert Berg, Gilad Saban, Gil Shabat
Large language model (LLM) agents are increasingly evaluated on their ability to use tools, plan multi-step tasks, coordinate with other agents, and operate over extended horizons. Reported benchmark gains often obscure recurring failure modes documented across otherwise unrelated evaluation efforts.
arXiv:2607. 05775v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly evaluated on their ability to use tools, plan multi-step tasks, coordinate with other agents, and operate over extended horizons.
By Wael Albayaydh, Rui Zhao, Ivan Flechais
arXiv:2607. 07729v1 Announce Type: cross Abstract: As foundation models grow in scale and diversity, coordinating multiple models into cooperative reasoning systems offers a path toward safer, more reliable AI.
By J. de Curt\`o, I. de Zarz\`a
arXiv:2608. 15389v1 Announce Type: new Abstract: LLM-based Text-to-SQL progress is reported across heterogeneous benchmarks, backbones, and inference protocols, making cross-system comparison fragile.
By Changruo Zhao, Zujun Peng, Yu Tian, Yuting Liu, Yiyun Su, Huiying Zhu, Luyan Zhang, Heming Zeng
Agentick is a unified benchmark for sequential decision‑making agents that evaluates RL, LLM, VLM, hybrid, and human agents on 37 procedurally generated tasks across six capability categories, four difficulty levels, and five observation modalities via a single Gymnasium‑compatible interface. It includes a Coding API, oracle reference policies, pre‑built SFT datasets, a composable agent harness, and a live leaderboard. An evaluation of 27 configurations and over 90,000 episodes shows no single approach dominates, with GPT‑5 mini leading overall, PPO excelling in planning and multi‑agent tasks, and the reasoning harness boosting LLM performance by 3‑10×, while ASCII observations outperform natural language.
By Roger Creus Castanyer, Pablo Samuel Castro, Glen Berseth
arXiv:2510. 15416v2 Announce Type: replace Abstract: We investigate a framework in which LoRA adapters are treated as callable tools that a base language model can dynamically select and invoke.
By Pavan C Shekar, Aswanth Krishnan
arXiv:2607. 05477v1 Announce Type: cross Abstract: Improving the task performance of Large Language Models (LLMs) is essential, yet scaling these models faces significant challenges such as diminishing returns and high costs.
By Lars Benedikt Kaesberg
arXiv:2607. 20499v1 Announce Type: new Abstract: Large Language Models generate plausible backend code, but a single-pass paradigm provides no guarantee of correctness or runtime reliability.
By Sai Deekshith Lekkala, Jothi Prabha Appadurai, Rohith Reddy Bellibatlu, Manpreet Singh
Loom is a generative consensus framework designed for real‑world root‑cause analysis (RCA) that combines open‑form hypotheses from modular heuristics with a lightweight large language model (LLM). It projects hypotheses into a continuous embedding space and uses an iterative centroid‑based reweighting algorithm to resolve conflicts, producing a single consensus that is then synthesized by one LLM call. On the OpenRCA benchmark, Loom achieves state‑of‑the‑art accuracy on Bank and Market‑2 while being significantly faster and more efficient than existing autonomous agents.