arXiv AI

Self-Organizing Agent Teams Learn to Reason Together

arXiv AI
Jul 23

Knowledge-Centric Self-Improvement

arXiv:2607. 19592v1 Announce Type: new Abstract: Self-improving AI systems typically treat the agent as the object that improves, by optimizing prompts, workflows, harnesses, or even the agent's own code.

By Xuefei Julie Wang, Lauren Hyoseo Yoon, Chengrui Qu, Amanda Zichang Wang, Atharva Sehgal, Eric Mazumdar, Yisong Yue
arXiv AI
Jul 29

Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm

arXiv:2607. 25446v1 Announce Type: new Abstract: Multi-agent frameworks built on large language models (LLMs) routinely entangle three logically distinct concerns: who is on the team (organization), how members align (coordination), and which algorithm fuses their work (collaboration protocol).

By Huan Chen, Xiang Song, Jian Jin, Pan Ren, Liang-Jie Zhang
arXiv AI
Sep 21

Scaling Discovery through Test-Time Communication

The paper demonstrates that test‑time communication among agents can significantly outperform independent parallel attempts on complex tasks. In experiments on the ARC‑AGI‑3 benchmark, a team of $k$ communicating agents matched the success rate of $4k$ independent agents, with the advantage growing as the team size increased. The study also shows that communication enables solving tasks that no single agent can solve, and that these benefits transfer to research‑oriented problems such as polyomino packing and MNIST classifier compression, where communicating agents surpassed prior best scores.

By Jongho Park, Vasilis Kontonis, Shivam Garg, Akshay Krishnamurthy, Dimitris Papailiopoulos
arXiv Computation and Language
Sep 23

Agensh: Scaling Organizational Intelligence to 1,024 Agents

Agensh is a new multi‑agent harness that eliminates a central orchestrator by letting workers self‑organize through a continuous cooperation loop. The system uses a shared workspace, message interface, and shared context to coordinate tasks, verify results, and merge progress asynchronously. Experiments on ProgramBench and pandoc show that scaling from 1 to 1,024 agents improves test‑pass rates by up to 49% relative, demonstrating that agent count is a viable scaling dimension for complex tasks.

By Zhihao Zhan, Ting Song, Li Dong, Shaohan Huang, Jianxun Lian, Yan Xia, Furu Wei
arXiv Computation and Language
Sep 23

CONCAT: Consensus- and Confidence-Driven Ad Hoc Teaming for Efficient LLM-Based Multi-Agent Systems

CONCAT is a training‑free framework that improves the efficiency of large language model (LLM) based multi‑agent systems by clustering agents according to their initial answers and selecting cluster leaders based on confidence. It uses a Theory‑of‑Mind‑inspired heuristic to predict collaboration benefits between leaders, then prunes communications to form an ad‑hoc network that reduces latency. Experiments on three LLMs and benchmarks show up to 2.02× higher accuracy/latency ratio than LLM‑Debate and a 50.1% latency reduction on Qwen2.5‑14B‑Instruct without task‑specific training.

By Ziyang Ma, Dingyi Zhang, Sichu Liang, Jiajia Chu, Pengfei Xia, Hui Zang, Deyu Zhou
arXiv AI
Aug 11

The Collaboration Gap: Exploration and Benchmarking of Open-World Agentic Cooperation

arXiv:2511. 02687v2 Announce Type: replace Abstract: The trajectory of AI development suggests that we will increasingly rely on agent-based systems powered by language models, composed of independently developed agents with different information, privileges, and tools.

By Tim R. Davidson, Adam Fourney, Saleema Amershi, Robert West, Eric Horvitz, Ece Kamar
arXiv AI
Jun 6

Humans' ALMANAC: A Human Collaboration Dataset of Action-Level Mental Model Annotations for Agent Collaboration

arXiv:2606. 06388v1 Announce Type: new Abstract: Recent advances in LLM agents have enabled complex cognitive capabilities, such as multi-step reasoning, planning, and tool use, that increasingly position these agents as human collaborators.

By Jiaju Chen, Yuxuan Lu, Jiayi Su, Chaoran Chen, Songlin Xiao, Zheng Zhang, Yun Wang, Yunyao Li, Jian Zhao, Tongshuang Wu, Toby Jia-Jun Li, Dakuo Wang, Bingsheng Yao
arXiv AI
Jul 16

Networked Intelligence: Active Shared Context Graphs for Human-AI Team Science

arXiv:2607. 13220v1 Announce Type: new Abstract: Most AI-for-science systems focus on scaling a single reasoning process through better models, larger context windows, long-horizon agentic execution, or digital co-scientists working with one principal user.

By Sutanay Choudhury, Jeffrey J. Czajka, Lummy M. O. Monteiro, Erin Bredeweg, Jason McDermott, Katherine Wolf, Alex Beliaev, Josh Elmore, Paul Piehowski, Kylee Tate, Yuqian Gao, Aivett Bilbao, Kelly Stratton, Scott Baker, Jaydeep P. Bardhan, Kristin Burnum Johnson, Chris Oehmen, Robert Rallo
arXiv AI
Sep 18

Communication and Verification in LLM Agents towards Collaboration under Information Asymmetry

The paper investigates how Large Language Model (LLM) agents can collaborate on a shared task under information asymmetry, using a table‑top version of Einstein Puzzles. It introduces a fine‑tuning‑plus‑verifier framework that equips agents with communication strategies and environmental verification signals. Results show that aligned communication is crucial for rule understanding and human trust, while a verifier improves task comprehension and promotes safer, interpretable collaboration.

By Run Peng, Ziqiao Ma, Amy Pang, Sikai Li, Zhang Xi-Jia, Yingzhuo Yu, Cristian-Paul Bara, Joyce Chai