arXiv Machine Learning By Kairui Yang, Xunkai Li, Kaixiang Zhang, Minghao An, Zekai Chen, Yuxuan Ba, Rong-Hua Li

OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems

Read the original on arXiv Machine Learning →

OpenMAS-GCom is a diagnostic benchmark designed to isolate the impact of communication structures, role assignments, and information flows in graph‑enhanced multi‑agent systems (G‑MAS). It evaluates systems by systematically modifying one component—such as rewiring communication edges, removing specialist or critic agents, or corrupting intermediate messages—while keeping tasks, models, prompts, and budget limits constant. The benchmark tests 17 configurations across 29 datasets in six domains, including 400 new G‑MAS‑Complex tasks that require agents to combine and reconcile information from multiple documents.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 29

Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm

arXiv:2607. 25446v1 Announce Type: new Abstract: Multi-agent frameworks built on large language models (LLMs) routinely entangle three logically distinct concerns: who is on the team (organization), how members align (coordination), and which algorithm fuses their work (collaboration protocol).

By Huan Chen, Xiang Song, Jian Jin, Pan Ren, Liang-Jie Zhang
arXiv AI
Jun 9

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents

arXiv:2606. 08340v1 Announce Type: new Abstract: As language models are increasingly deployed as autonomous agents, they must coordinate with others over long horizons in open-ended interactive tasks.

By Kale-ab Abebe Tessera, Andras Szecsenyi, Cameron Barker, Alexander Rutherford, Davide Paglieri, Aidan Scannell, Henry Gouk, Elliot J. Crowley, Tim Rockt\"aschel, Amos Storkey
arXiv AI
Sep 21

DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement

arXiv:2609.21423v1 Announce Type: new Abstract: Online agent deployments produce abundant execution traces, while task-specific verification and expert annotation are costly to scale. We study how to...

By Siyuan Liu (Fudan University, Meituan Longcat Team), Fan Yu (Fudan University, Meituan Longcat Team), Dongyu Ru (Meituan Longcat Team), Yizhu Liu (Meituan Longcat Team), Yifan Yang (Meituan Longcat Team), Xuezhi Cao (Meituan Longcat Team), Xunliang Cai (Meituan Longcat Team), Yixin Cao (Fudan University)