arXiv AI

Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm

arXiv:2607. 25446v1 Announce Type: new Abstract: Multi-agent frameworks built on large language models (LLMs) routinely entangle three logically distinct concerns: who is on the team (organization), how members align (coordination), and which algorithm fuses their work (collaboration protocol).

arXiv Computation and Language
3d ago

You're Hired: Strategic Model Selection for LLM Collaboration

arXiv:2609.38816v1 Announce Type: new Abstract: While multi-agent and model collaboration algorithms gain traction to combine the strengths of diverse Large Language Models (LLMs), existing systems r...

By Zongwan Cao, Ziyuan Yang, Shangbin Feng, Michael Duan, Skyler Hallinan, Bingbing Wen, Lucy Lu Wang, Yulia Tsvetkov
arXiv AI
3d ago

CollabFlow: Recursive Self-Improvement of Agent Collaboration

CollabFlow introduces a recursive self‑improvement framework for multi‑agent collaboration in large language model systems. It trains a Collab‑Director to assemble teams of agents, uses a frozen executor to run them, and retrains the director each round based on outcomes. The system incorporates evidence‑conditioned communication protocols within collaboration graphs and a Collaborative Trajectory Balance objective to maintain diverse high‑performing teams across rounds, achieving superior performance on twelve datasets.

By Xiao Huang, Mingda Zhang, Junming Zhang, Qiang Huang, Hanwen Zhang, Yue Dai, Zijia Wang, Xiaoying Tang
arXiv AI
Jun 9

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents

arXiv:2606. 08340v1 Announce Type: new Abstract: As language models are increasingly deployed as autonomous agents, they must coordinate with others over long horizons in open-ended interactive tasks.

By Kale-ab Abebe Tessera, Andras Szecsenyi, Cameron Barker, Alexander Rutherford, Davide Paglieri, Aidan Scannell, Henry Gouk, Elliot J. Crowley, Tim Rockt\"aschel, Amos Storkey
arXiv Machine Learning
Sep 21

OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems

OpenMAS-GCom is a diagnostic benchmark designed to isolate the impact of communication structures, role assignments, and information flows in graph‑enhanced multi‑agent systems (G‑MAS). It evaluates systems by systematically modifying one component—such as rewiring communication edges, removing specialist or critic agents, or corrupting intermediate messages—while keeping tasks, models, prompts, and budget limits constant. The benchmark tests 17 configurations across 29 datasets in six domains, including 400 new G‑MAS‑Complex tasks that require agents to combine and reconcile information from multiple documents.

By Kairui Yang, Xunkai Li, Kaixiang Zhang, Minghao An, Zekai Chen, Yuxuan Ba, Rong-Hua Li
arXiv AI
2d ago

Worse Together: How Performance Breaks Down in Multi-User Multi-Agent Teams

The paper investigates how coordination among AI agents serving different users degrades performance compared to a single coordinating agent. Across five advanced models and 77 scenarios in four shared-resource environments—API key budgets, clinic calendars, personal assistant bookings, and merge queues—the study finds that multi‑agent teams consistently underperform, sometimes collapsing entirely, and that even with communication channels coordination overhead remains significant. The authors identify specific failure modes such as stalling, action overriding, and claim fabrication, and propose environment‑specific mitigations like team leads and procedural instructions, while releasing the MAMUBench benchmark for future research.

By Sahan Paliskara, Nattaput Namchittai, Andrew Lampinen
arXiv AI
Sep 11

MOSAIC: A Universal Agent-Level Interface for Cross-Paradigm Agent Mixing and Human-AI Collaboration

MOSAIC is an open‑source platform that allows agents from different decision‑making paradigms—such as reinforcement learning policies, large language models, vision‑language models, and human operators—to operate together in shared reinforcement learning environments. It achieves this through an IPC‑based worker protocol that isolates each agent’s training and inference logic, an operator abstraction that maps any agent to a minimal universal interface, and a deterministic evaluation framework offering both manual lock‑step and automated script modes for reproducible experiments.

By Abdulhamid M. Mousa, Jinhui Pang, Rakhmonberdi Khajiev, Jalaledin M. Azzabi, Abdulkarim M. Mousa, Peng Yong, Yunusa Haruna, Ming Liu