arXiv AI By Xiao Huang, Mingda Zhang, Junming Zhang, Qiang Huang, Hanwen Zhang, Yue Dai, Zijia Wang, Xiaoying Tang

CollabFlow: Recursive Self-Improvement of Agent Collaboration

Read the original on arXiv AI →

CollabFlow introduces a recursive self‑improvement framework for multi‑agent collaboration in large language model systems. It trains a Collab‑Director to assemble teams of agents, uses a frozen executor to run them, and retrains the director each round based on outcomes. The system incorporates evidence‑conditioned communication protocols within collaboration graphs and a Collaborative Trajectory Balance objective to maintain diverse high‑performing teams across rounds, achieving superior performance on twelve datasets.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
23h ago

EvoSteer: Online Self-Evolving Graph Orchestration via Reference-Anchored Credit Assignment

EvoSteer introduces an online self‑evolving graph orchestration framework that continuously builds and repairs a team of tool‑using agents during execution. It employs Anchored Trajectory Balance (AnchorTB), a regression‑style loss that assigns credit to each orchestration action by comparing subtrajectories to a frozen reference, and Validated Skill Admission, which tests candidate skills before promotion. Experiments on twelve datasets demonstrate that EvoSteer outperforms existing baselines in question answering, mathematical reasoning, code generation, and interactive decision making.

By Mingda Zhang, Hanwen Zhang, Qiang Huang, Zijia Wang, Pengfei Guo, Yuchen Zhang, Jionghao Zhu, Xiaoying Tang
arXiv Computation and Language
Sep 23

CONCAT: Consensus- and Confidence-Driven Ad Hoc Teaming for Efficient LLM-Based Multi-Agent Systems

CONCAT is a training‑free framework that improves the efficiency of large language model (LLM) based multi‑agent systems by clustering agents according to their initial answers and selecting cluster leaders based on confidence. It uses a Theory‑of‑Mind‑inspired heuristic to predict collaboration benefits between leaders, then prunes communications to form an ad‑hoc network that reduces latency. Experiments on three LLMs and benchmarks show up to 2.02× higher accuracy/latency ratio than LLM‑Debate and a 50.1% latency reduction on Qwen2.5‑14B‑Instruct without task‑specific training.

By Ziyang Ma, Dingyi Zhang, Sichu Liang, Jiajia Chu, Pengfei Xia, Hui Zang, Deyu Zhou
arXiv AI
Jul 23

In-the-Flow Agentic System Optimization for Effective Planning and Tool Use

arXiv:2510. 05592v2 Announce Type: replace Abstract: Outcome-driven reinforcement learning has advanced reasoning in large language models (LLMs), but prevailing tool-augmented approaches train a single, monolithic policy that interleaves thoughts and tool calls under full context; this scales poorly with long horizons and diverse tools and generalizes weakly to new scenarios.

By Zhuofeng Li, Haoxiang Zhang, Seungju Han, Sheng Liu, Jianwen Xie, Yu Zhang, Yejin Choi, James Zou, Pan Lu
Hugging Face Trending Papers
Jun 18

Connect the Dots: Training LLMs for Long-Lifecycle Agents with Cross-Domain Generalization Via Reinforcement Learning

This work presents a general framework for training large language models (LLMs) to "Connect the Dots" (CoD), a meta-capability required by long-lifecycle agents: as an LLM-based AI agent gets deployed in an environment, it solves a long sequence of tasks while continuously exploring the environment, learning from its own experiences, and iteratively self-updating its context about the environment, thereby achieving progressively better performance on future tasks conditioned on the updated context. Major components of the CoD framework include: (1) algorithm design and infrastructure for end-to-end reinforcement learning (RL) with long rollout sequences interleaving solve-task and update-context episodes; (2) tasks and environments for incentivizing and eliciting the targeted meta-capability in LLMs during training, as well as for faithfully measuring progress during evaluation.

arXiv AI
Jul 29

Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm

arXiv:2607. 25446v1 Announce Type: new Abstract: Multi-agent frameworks built on large language models (LLMs) routinely entangle three logically distinct concerns: who is on the team (organization), how members align (coordination), and which algorithm fuses their work (collaboration protocol).

By Huan Chen, Xiang Song, Jian Jin, Pan Ren, Liang-Jie Zhang
arXiv AI
Sep 11

ROTATE: Regret-driven Open-ended Training for Ad Hoc Teamwork

The paper introduces ROTATE, a regret-driven open‑ended training framework that jointly improves an Ad Hoc Teamwork (AHT) agent and an adversarial teammate generator. Unlike traditional two‑stage pipelines, ROTATE alternates between enhancing the agent and generating teammates that specifically probe its collaboration weaknesses. Experiments on Overcooked and Level‑Based Foraging show that ROTATE outperforms existing baselines on unseen teammates, setting a new benchmark for robust, generalizable teamwork.

By Caroline Wang, Arrasy Rahman, Benjamin Nativi, Jiaxun Cui, Yoonchang Sung, Peter Stone