The paper investigates challenges in open‑source large‑language‑model (LLM) based multi‑agent systems (MAS). By analyzing 944 issues extracted from 21 projects, it finds that orchestration and execution problems are most common, with workflow, tool integration, and memory issues as primary causes. The predominant remedy identified is optimizing workflow, and the study offers empirically grounded implications for improving orchestration, tool integration, and memory mechanisms in LLM‑based MAS.
By Asad Ur Rehman, Syed Mohammad Kashif, Ruiyin Li, Peng Liang, Zengyang Li, Arif Ali Khan
arXiv:2607. 13705v1 Announce Type: new Abstract: As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical.
By Zichen Ding, Jiaye Ge, Shufan Jiang, Kai Chen, Mo Li, Qingqiu Li, Zehao Li, Zonglin Li, Tiaohao Liang, Shudong Liu, Zerun Ma, Zixing Shang, Wenhui Tian, Zun Wang, Liwei Wu, Zhenyu Wu, Jun Xu, Bowen Yang, Dingbo Yuan, Qi Zhang, Songyang Zhang, Peiheng Zhou, Dongsheng Zhu
arXiv:2602. 22480v4 Announce Type: replace Abstract: An important emerging application of coding agents is agent harness optimization: the iterative improvement of a target agent by editing and evaluating its code.
By Varun Ursekar, Apaar Shanker, Veronica Chatrath, Yuan Xue, Samuel Marc Denton
As Large Language Models (LLMs) evolve into autonomous agents, the need for unified evaluation infrastructure becomes critical. However, current evaluation pipelines remain highly fragmented and tightly coupled, hindering reproducibility and causing redundant engineering.
arXiv:2505. 11765v5 Announce Type: replace-cross Abstract: Agents powered by advanced large language models (LLMs) have demonstrated impressive capabilities across diverse complex applications.
By Shijun Li, Hilaf Hasson, Joydeep Ghosh
arXiv:2605. 27898v2 Announce Type: replace Abstract: As LLMs are increasingly deployed as agents, reliable assessment of their agentic capabilities has become essential.
By Pengyu Zhu, Lijun Li, Yaxing Lyu, Qianxin Luo, Jingyi Yang, Yi Liu, Tingfeng Hui, Xinyu Yuan, Li Sun, Sen Su, Jing Shao
The paper critiques current evaluations of efficiency methods for large language model–based multi‑agent systems, arguing that reported gains are often inflated by method‑specific prompts and starting topologies. It introduces a controlled, MAS‑demanding diagnostic benchmark that standardizes the backbone model, agent registry, and runtime, and systematically varies topology, scale, depth, and tool use. The authors find that many claimed efficiency improvements are setup‑dependent, sometimes stemming from structural collapse or random pruning rather than genuine, robust gains.
By Jiamu Zhang, Lingxi Zhang, Pengjun Lu, Qiyue Zhang, Yu-Neng Chuang, Zhengchen Li, Shuai Xu, Vipin Chaudhary, Hanjie Chen
arXiv:2606. 13608v1 Announce Type: new Abstract: Agent systems are advancing quickly across domains, but their evaluation remains fragmented.
By Xiaoyuan Liu, Jianhong Tu, Yuqi Chen, Siyuan Xie, Sihan Ren, Tianneng Shi, Gal Gantar, Evan Sandoval, Donghyun Lee, Daniel Miao, Peter J. Gilbert, Nick Hynes, Mauro Staver, Warren He, David Marn, Andrew Low, Xi Zhang, Elron Bandel, Michal Shmueli-Scheuer, Siva Reddy, Alexandre Drouin, Alexandre Lacoste, Ramayya Krishnan, Elham Tabassi, Yu Su, Victor Barres, Chenguang Wang, Wenbo Guo, Dawn Song
OpenCollab is a multi‑agent coding framework that unifies organization design, enforces experimental control on a shared runtime, and tracks execution via fine‑grained event streams. It introduces the metric Adherence to measure whether the declared organization is actually realized, showing that small configuration changes can shift adherence from 47.2% to 97.2%. Experiments demonstrate that a two‑coder workflow built on OpenCollab achieves new state‑of‑the‑art performance against mainstream harnesses while using the fewest tokens, and that a well‑designed organization can outperform strong existing harnesses.
By Chun-Wah Hsu, Kai Gong, Yu Wu, Xianhe Chen, Mengyang Liu, Jie Li, Hanyu Li, Zhixuan Liu, Naisheng Tang, Jiaying Chi, Ziheng Fan, Xuning He, Xiaokang Yang, Xue Jiang, Yihong Dong
Multi-agent Systems (MAS) combine multiple model outputs to solve complex reasoning tasks. However, despite rapid growth of available open-source models, there is limited research on how to select opt...
arXiv:2606. 05670v1 Announce Type: new Abstract: Does adding more agents help an LLM workflow once compared systems share the same benchmark loader, tool access, answer contract, usage accounting, and trajectory logging?
By Yuhang Fu, Ruishan Fang, Jiaqi Shao, Huiyu Zheng, Zhengtao Zhu, Bing Luo, Tao Lin
UniACE is a unified framework that standardizes the evaluation of large language model (LLM) agents by representing each benchmark as an instruction–tool–environment triplet and running models through a shared, task‑agnostic harness in isolated runtimes. It preserves native success criteria, offers an offline mode for dynamic‑resource tasks, and standardizes efficiency metrics, execution records, and failure attribution. Applying UniACE to 7 benchmarks across 24 domains and 15 models revealed significant score shifts, ranking reversals, and sensitivity to evidence representation, highlighting the impact of evaluation configuration on reported agent performance.
By Pengyu Zhu, Lijun Li, Yaxing Lyu, Qianxin Luo, Jingyi Yang, Yi Liu, Tingfeng Hui, Xinyu Yuan, Li Sun, Sen Su, Jing Shao