arXiv Machine Learning

OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems

OpenMAS-GCom is a diagnostic benchmark designed to isolate the impact of communication structures, role assignments, and information flows in graph‑enhanced multi‑agent systems (G‑MAS). It evaluates systems by systematically modifying one component—such as rewiring communication edges, removing specialist or critic agents, or corrupting intermediate messages—while keeping tasks, models, prompts, and budget limits constant. The benchmark tests 17 configurations across 29 datasets in six domains, including 400 new G‑MAS‑Complex tasks that require agents to combine and reconcile information from multiple documents.

arXiv AI
Jul 29

Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm

arXiv:2607. 25446v1 Announce Type: new Abstract: Multi-agent frameworks built on large language models (LLMs) routinely entangle three logically distinct concerns: who is on the team (organization), how members align (coordination), and which algorithm fuses their work (collaboration protocol).

By Huan Chen, Xiang Song, Jian Jin, Pan Ren, Liang-Jie Zhang
arXiv AI
Jun 9

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents

arXiv:2606. 08340v1 Announce Type: new Abstract: As language models are increasingly deployed as autonomous agents, they must coordinate with others over long horizons in open-ended interactive tasks.

By Kale-ab Abebe Tessera, Andras Szecsenyi, Cameron Barker, Alexander Rutherford, Davide Paglieri, Aidan Scannell, Henry Gouk, Elliot J. Crowley, Tim Rockt\"aschel, Amos Storkey
arXiv AI
Sep 21

DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement

arXiv:2609.21423v1 Announce Type: new Abstract: Online agent deployments produce abundant execution traces, while task-specific verification and expert annotation are costly to scale. We study how to...

By Siyuan Liu (Fudan University, Meituan Longcat Team), Fan Yu (Fudan University, Meituan Longcat Team), Dongyu Ru (Meituan Longcat Team), Yizhu Liu (Meituan Longcat Team), Yifan Yang (Meituan Longcat Team), Xuezhi Cao (Meituan Longcat Team), Xunliang Cai (Meituan Longcat Team), Yixin Cao (Fudan University)
arXiv AI
Sep 18

Rethinking Multi-Agent Collaboration: When More Is Less

The paper examines when multi‑agent collaboration is beneficial versus single‑agent approaches. It finds that collaboration yields systematic advantages mainly in long‑horizon tasks with sparse dependencies, while single agents perform better in tightly coupled, sequential workflows. The authors introduce SAIGE, a lightweight multi‑agent mechanism that models collaboration as a dynamically evolving graph, and show that it balances context efficiency and task performance without always improving outcomes as more agents are added.

By Yishuo Yuan, Yibo Wu, Yihan Zhang, Minyuan Sun, Shenliang Li, Xinkai Ma, Yifan Li, Jiaheng Liu
arXiv AI
Aug 11

The Collaboration Gap: Exploration and Benchmarking of Open-World Agentic Cooperation

arXiv:2511. 02687v2 Announce Type: replace Abstract: The trajectory of AI development suggests that we will increasingly rely on agent-based systems powered by language models, composed of independently developed agents with different information, privileges, and tools.

By Tim R. Davidson, Adam Fourney, Saleema Amershi, Robert West, Eric Horvitz, Ece Kamar
arXiv AI
Sep 2

UniACE: A Unified Framework for Evaluating LLM Agentic Capabilities

UniACE is a unified framework that standardizes the evaluation of large language model (LLM) agents by representing each benchmark as an instruction–tool–environment triplet and running models through a shared, task‑agnostic harness in isolated runtimes. It preserves native success criteria, offers an offline mode for dynamic‑resource tasks, and standardizes efficiency metrics, execution records, and failure attribution. Applying UniACE to 7 benchmarks across 24 domains and 15 models revealed significant score shifts, ranking reversals, and sensitivity to evidence representation, highlighting the impact of evaluation configuration on reported agent performance.

By Pengyu Zhu, Lijun Li, Yaxing Lyu, Qianxin Luo, Jingyi Yang, Yi Liu, Tingfeng Hui, Xinyu Yuan, Li Sun, Sen Su, Jing Shao