arXiv AI By Huan Chen, Xiang Song, Jian Jin, Pan Ren, Liang-Jie Zhang

Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm

Read the original on arXiv AI →

arXiv:2607. 25446v1 Announce Type: new Abstract: Multi-agent frameworks built on large language models (LLMs) routinely entangle three logically distinct concerns: who is on the team (organization), how members align (coordination), and which algorithm fuses their work (collaboration protocol).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
3d ago

You're Hired: Strategic Model Selection for LLM Collaboration

arXiv:2609.38816v1 Announce Type: new Abstract: While multi-agent and model collaboration algorithms gain traction to combine the strengths of diverse Large Language Models (LLMs), existing systems r...

By Zongwan Cao, Ziyuan Yang, Shangbin Feng, Michael Duan, Skyler Hallinan, Bingbing Wen, Lucy Lu Wang, Yulia Tsvetkov
arXiv AI
3d ago

CollabFlow: Recursive Self-Improvement of Agent Collaboration

CollabFlow introduces a recursive self‑improvement framework for multi‑agent collaboration in large language model systems. It trains a Collab‑Director to assemble teams of agents, uses a frozen executor to run them, and retrains the director each round based on outcomes. The system incorporates evidence‑conditioned communication protocols within collaboration graphs and a Collaborative Trajectory Balance objective to maintain diverse high‑performing teams across rounds, achieving superior performance on twelve datasets.

By Xiao Huang, Mingda Zhang, Junming Zhang, Qiang Huang, Hanwen Zhang, Yue Dai, Zijia Wang, Xiaoying Tang
arXiv AI
Jun 9

Benchmarking Open-Ended Multi-Agent Coordination in Language Agents

arXiv:2606. 08340v1 Announce Type: new Abstract: As language models are increasingly deployed as autonomous agents, they must coordinate with others over long horizons in open-ended interactive tasks.

By Kale-ab Abebe Tessera, Andras Szecsenyi, Cameron Barker, Alexander Rutherford, Davide Paglieri, Aidan Scannell, Henry Gouk, Elliot J. Crowley, Tim Rockt\"aschel, Amos Storkey
arXiv Machine Learning
Sep 21

OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems

OpenMAS-GCom is a diagnostic benchmark designed to isolate the impact of communication structures, role assignments, and information flows in graph‑enhanced multi‑agent systems (G‑MAS). It evaluates systems by systematically modifying one component—such as rewiring communication edges, removing specialist or critic agents, or corrupting intermediate messages—while keeping tasks, models, prompts, and budget limits constant. The benchmark tests 17 configurations across 29 datasets in six domains, including 400 new G‑MAS‑Complex tasks that require agents to combine and reconcile information from multiple documents.

By Kairui Yang, Xunkai Li, Kaixiang Zhang, Minghao An, Zekai Chen, Yuxuan Ba, Rong-Hua Li