arXiv Machine Learning

Assembling the CREW: A Collaborative Multi-agent Reinforcement Learning Framework for Automated Related Work Generation

The paper introduces CREW, a collaborative multi‑agent reinforcement learning framework that automates the generation of the Related Work Section in research papers. Unlike previous methods that follow a fixed workflow, CREW allows large language model agents to dynamically select actions—Retrieve, Disseminate, Compose, and Critique—guided by a policy trained with Independent Proximal Policy Optimization. Experiments on a standard benchmark show that CREW improves output quality and reduces token usage compared to strong baselines.

arXiv AI
Aug 28

DeepPlanner: Scaling Planning Capability for Deep Research Agents via Advantage Shaping

DeepPlanner is an end-to-end reinforcement learning framework designed to enhance the planning capabilities of deep research agents. It introduces an entropy-based advantage shaping mechanism that allocates larger updates to high-entropy planning tokens and selectively upweights sample-level advantages during planning-intensive rollouts. Experiments on seven deep research benchmarks show that DeepPlanner improves planning quality and achieves state‑of‑the‑art results with a lower training budget.

By Wei Fan, Wenlin Yao, Zheng Li, Feng Yao, Xin Liu, Liang Qiu, Qingyu Yin, Yangqiu Song, Bing Yin
Hugging Face Trending Papers
Jul 29

SciDataSailor: Deep Scientific Data Exploring

Scientific datasets are commonly organized as hierarchical repositories containing heterogeneous and interdependent files, making their inspection, integration, and analysis labor-intensive and reliant on domain expertise. Although large language model (LLM) agents have advanced substantially in planning, reasoning, and tool use, existing research has largely overlooked their ability to interact with real scientific data assets through executable environments.

arXiv AI
Sep 4

LDC: Learning to Generate Research Idea with Dynamic Control

The paper introduces LDC, a framework that learns to generate research ideas with dynamic control. It combines supervised fine‑tuning on paper‑idea pairs with controllable reinforcement learning that optimizes novelty, feasibility, and effectiveness. During inference, sentence‑level controllers steer the generation process to balance these dimensions.

By Ruochen Li, Liqiang Jing, Chi Han, Jiawei Zhou, Xinya Du
arXiv AI
Aug 26

Efficient LLM Collaboration via Planning

arXiv:2506.11578v5 Announce Type: replace Abstract: Recently, large language models (LLMs) have demonstrated strong performance, ranging from simple to complex tasks. However, while large models achi...

By Byeongchan Lee, Jonghoon Lee, Dongyoung Kim, Jaehyung Kim, Kyungjoon Park, Dongjun Lee, Jinwoo Shin
arXiv AI
Aug 19

PlanPO: Group Planning-Aware Policy Optimization for Multi-Turn Agentic LLMs

PlanPO introduces a group planning-aware policy optimization method for multi-turn agentic large language models, addressing the issue of advantage collapse caused by treating all successful trajectories equally. By incorporating coarse-to-fine advantage signals that reflect differences in trajectory and turn lengths, PlanPO encourages agents to learn generalizable planning and generation behaviors. Experiments show a 27.2% average improvement over GRPO on benchmarks such as ALFWorld, WebShop, and SciWorld, with minimal extra training cost.

By Dayang Liang, Liyuan He, Xuan Feng, Shuxin Li, Bo An, Yunlong Liu
arXiv AI
Sep 18

Rethinking Multi-Agent Collaboration: When More Is Less

The paper examines when multi‑agent collaboration is beneficial versus single‑agent approaches. It finds that collaboration yields systematic advantages mainly in long‑horizon tasks with sparse dependencies, while single agents perform better in tightly coupled, sequential workflows. The authors introduce SAIGE, a lightweight multi‑agent mechanism that models collaboration as a dynamically evolving graph, and show that it balances context efficiency and task performance without always improving outcomes as more agents are added.

By Yishuo Yuan, Yibo Wu, Yihan Zhang, Minyuan Sun, Shenliang Li, Xinkai Ma, Yifan Li, Jiaheng Liu
arXiv AI
Jul 23

In-the-Flow Agentic System Optimization for Effective Planning and Tool Use

arXiv:2510. 05592v2 Announce Type: replace Abstract: Outcome-driven reinforcement learning has advanced reasoning in large language models (LLMs), but prevailing tool-augmented approaches train a single, monolithic policy that interleaves thoughts and tool calls under full context; this scales poorly with long horizons and diverse tools and generalizes weakly to new scenarios.

By Zhuofeng Li, Haoxiang Zhang, Seungju Han, Sheng Liu, Jianwen Xie, Yu Zhang, Yejin Choi, James Zou, Pan Lu