arXiv AI

LDC: Learning to Generate Research Idea with Dynamic Control

The paper introduces LDC, a framework that learns to generate research ideas with dynamic control. It combines supervised fine‑tuning on paper‑idea pairs with controllable reinforcement learning that optimizes novelty, feasibility, and effectiveness. During inference, sentence‑level controllers steer the generation process to balance these dimensions.

arXiv AI
5d ago

Learning to Ideate for Scientific Impact

The paper "Learning to Ideate for Scientific Impact" explores using delayed signals of scientific uptake—specifically citation-normalized impact—as feedback to steer large language models toward generating high‑impact research ideas. The authors build a dataset of over 100,000 computer science papers, train a reward model to predict citation impact from goal‑idea pairs, and align an idea generator via supervised fine‑tuning and reinforcement learning. Evaluation with a reference‑grounded protocol shows that the RL‑tuned model consistently produces ideas with higher estimated impact than baseline models.

By Shubham Kale, Aniketh Garikaparthi, Manasi Patwardhan
arXiv Machine Learning
Sep 15

Assembling the CREW: A Collaborative Multi-agent Reinforcement Learning Framework for Automated Related Work Generation

The paper introduces CREW, a collaborative multi‑agent reinforcement learning framework that automates the generation of the Related Work Section in research papers. Unlike previous methods that follow a fixed workflow, CREW allows large language model agents to dynamically select actions—Retrieve, Disseminate, Compose, and Critique—guided by a policy trained with Independent Proximal Policy Optimization. Experiments on a standard benchmark show that CREW improves output quality and reduces token usage compared to strong baselines.

By Hai-Dang Dang, Bao-Yen Pham, Bao Nguyen, Tran Thi Huong, Huynh Thi Thanh Binh
arXiv AI
Sep 17

PrimeScientist: Strategic Allocation of Research Effort in Autonomous Research

PrimeScientist is a system that jointly selects research directions and allocates resources for autonomous research agents. It models the problem as a sequential decision task, using an executable plan tree to track competing plans and an adaptive MCTS-based policy to balance exploration and exploitation based on remaining resources and experimental feedback. Experiments on AI research, systems, code optimization, and machine learning engineering show that PrimeScientist improves average reward by 10.3% while reducing research attempts by 50.6% compared to AutoResearch under the same budget.

By Xinle Yu, Fan Bai, Kaiser Sun, Hengshuo Miao, Abhay Anand, Zhongyan Luo, Kun Zhou, Zhen Wang
arXiv Computation and Language
Sep 1

PaperGym: Rubric-Centered Evolution for Research-Plan Generation

PaperGym is a framework that transforms each research paper into a training environment for AI research planning, using rubrics extracted from the paper’s method and experiments as a critic. It synthesizes research questions from the goal and background, and derives evaluation criteria from the method and experiments, reducing criterion leakage to 3.7%. Experiments with Qwen models show that training with PaperGym’s rubric improves benchmark performance by up to 5.6 points and outperforms existing datasets and fine‑tuning baselines.

By Yuhan Wang, Zhengxi Lu, Yuchen Yan, Kaitao Song, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen
Hugging Face Trending Papers
Jul 29

SciDataSailor: Deep Scientific Data Exploring

Scientific datasets are commonly organized as hierarchical repositories containing heterogeneous and interdependent files, making their inspection, integration, and analysis labor-intensive and reliant on domain expertise. Although large language model (LLM) agents have advanced substantially in planning, reasoning, and tool use, existing research has largely overlooked their ability to interact with real scientific data assets through executable environments.

Hugging Face Trending Papers
Jun 3

Self-Evolving Deep Research via Joint Generation and Evaluation

Large Language Models (LLMs) have become increasingly adopted in daily applications, with deep research standing out as a particularly important capability. Unlike traditional question-answering (QA) tasks, deep research report generation lacks definitive ground-truth, making reward design inherently unverifiable and limiting effective reinforcement learning.

arXiv AI
Aug 28

DeepPlanner: Scaling Planning Capability for Deep Research Agents via Advantage Shaping

DeepPlanner is an end-to-end reinforcement learning framework designed to enhance the planning capabilities of deep research agents. It introduces an entropy-based advantage shaping mechanism that allocates larger updates to high-entropy planning tokens and selectively upweights sample-level advantages during planning-intensive rollouts. Experiments on seven deep research benchmarks show that DeepPlanner improves planning quality and achieves state‑of‑the‑art results with a lower training budget.

By Wei Fan, Wenlin Yao, Zheng Li, Feng Yao, Xin Liu, Liang Qiu, Qingyu Yin, Yangqiu Song, Bing Yin
arXiv AI
Sep 7

Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving

The paper investigates how the diversity of solutions produced by large language models (LLMs) for a single problem correlates with their problem‑solving performance. It finds that higher solution divergence is linked to better outcomes across various models and proposes using this metric to enhance supervised fine‑tuning and reinforcement learning. Experiments on three problem domains show that incorporating solution divergence consistently raises success rates, indicating its potential as a simple yet effective tool for LLM training and evaluation.

By Hang Li, Kaiqi Yang, Yucheng Chu, Hui Liu, Jiliang Tang
arXiv AI
Aug 19

AutoResearch: Insight In, Hallucination Out

AutoResearch is a two‑stage autonomous research system that links Idea Generation with Idea Execution. In the generation phase it blends new research signals with existing domain knowledge, identifies transferable mechanistic insights, and produces grounded, testable research plans through multi‑model generation and cross‑review. The execution phase then decomposes these plans into experiments, iteratively implements and diagnoses them, and uses independent evidence‑based review to accept or revise conclusions, thereby turning ideas into measurable progress while minimizing hallucinations.

By Yiming Ren, Xiang Liu, Qumeng Sun, Xiao Zhang, Jiahao Li, Haoyang Zhang, Junjie Wang