The paper "Learning to Ideate for Scientific Impact" explores using delayed signals of scientific uptake—specifically citation-normalized impact—as feedback to steer large language models toward generating high‑impact research ideas. The authors build a dataset of over 100,000 computer science papers, train a reward model to predict citation impact from goal‑idea pairs, and align an idea generator via supervised fine‑tuning and reinforcement learning. Evaluation with a reference‑grounded protocol shows that the RL‑tuned model consistently produces ideas with higher estimated impact than baseline models.
By Shubham Kale, Aniketh Garikaparthi, Manasi Patwardhan
The paper introduces CREW, a collaborative multi‑agent reinforcement learning framework that automates the generation of the Related Work Section in research papers. Unlike previous methods that follow a fixed workflow, CREW allows large language model agents to dynamically select actions—Retrieve, Disseminate, Compose, and Critique—guided by a policy trained with Independent Proximal Policy Optimization. Experiments on a standard benchmark show that CREW improves output quality and reduces token usage compared to strong baselines.
By Hai-Dang Dang, Bao-Yen Pham, Bao Nguyen, Tran Thi Huong, Huynh Thi Thanh Binh
arXiv:2608.30109v1 Announce Type: new
Abstract: Large Language Models (LLMs) trained on extensive scientific research are increasingly integrated as assistants for scientific discovery. However, most...
By Shrinidhi Kumbhar Santosh Mashetty Divij Handa Kevin Coutinho, Siddharth Sambhaji Ghule, Chitta Baral
arXiv:2606. 04507v1 Announce Type: cross Abstract: Large Language Models (LLMs) have become increasingly adopted in daily applications, with deep research standing out as a particularly important capability.
By Han Zhu, Chengkun Cai, Yuanfeng Song, Xing Chen, Sirui Han, Yike Guo
PrimeScientist is a system that jointly selects research directions and allocates resources for autonomous research agents. It models the problem as a sequential decision task, using an executable plan tree to track competing plans and an adaptive MCTS-based policy to balance exploration and exploitation based on remaining resources and experimental feedback. Experiments on AI research, systems, code optimization, and machine learning engineering show that PrimeScientist improves average reward by 10.3% while reducing research attempts by 50.6% compared to AutoResearch under the same budget.
By Xinle Yu, Fan Bai, Kaiser Sun, Hengshuo Miao, Abhay Anand, Zhongyan Luo, Kun Zhou, Zhen Wang
PaperGym is a framework that transforms each research paper into a training environment for AI research planning, using rubrics extracted from the paper’s method and experiments as a critic. It synthesizes research questions from the goal and background, and derives evaluation criteria from the method and experiments, reducing criterion leakage to 3.7%. Experiments with Qwen models show that training with PaperGym’s rubric improves benchmark performance by up to 5.6 points and outperforms existing datasets and fine‑tuning baselines.
By Yuhan Wang, Zhengxi Lu, Yuchen Yan, Kaitao Song, Wenqi Zhang, Weiming Lu, Jun Xiao, Yueting Zhuang, Yongliang Shen
Scientific datasets are commonly organized as hierarchical repositories containing heterogeneous and interdependent files, making their inspection, integration, and analysis labor-intensive and reliant on domain expertise. Although large language model (LLM) agents have advanced substantially in planning, reasoning, and tool use, existing research has largely overlooked their ability to interact with real scientific data assets through executable environments.
Large Language Models (LLMs) have become increasingly adopted in daily applications, with deep research standing out as a particularly important capability. Unlike traditional question-answering (QA) tasks, deep research report generation lacks definitive ground-truth, making reward design inherently unverifiable and limiting effective reinforcement learning.
DeepPlanner is an end-to-end reinforcement learning framework designed to enhance the planning capabilities of deep research agents. It introduces an entropy-based advantage shaping mechanism that allocates larger updates to high-entropy planning tokens and selectively upweights sample-level advantages during planning-intensive rollouts. Experiments on seven deep research benchmarks show that DeepPlanner improves planning quality and achieves state‑of‑the‑art results with a lower training budget.
By Wei Fan, Wenlin Yao, Zheng Li, Feng Yao, Xin Liu, Liang Qiu, Qingyu Yin, Yangqiu Song, Bing Yin
The paper investigates how the diversity of solutions produced by large language models (LLMs) for a single problem correlates with their problem‑solving performance. It finds that higher solution divergence is linked to better outcomes across various models and proposes using this metric to enhance supervised fine‑tuning and reinforcement learning. Experiments on three problem domains show that incorporating solution divergence consistently raises success rates, indicating its potential as a simple yet effective tool for LLM training and evaluation.
By Hang Li, Kaiqi Yang, Yucheng Chu, Hui Liu, Jiliang Tang
AutoResearch is a two‑stage autonomous research system that links Idea Generation with Idea Execution. In the generation phase it blends new research signals with existing domain knowledge, identifies transferable mechanistic insights, and produces grounded, testable research plans through multi‑model generation and cross‑review. The execution phase then decomposes these plans into experiments, iteratively implements and diagnoses them, and uses independent evidence‑based review to accept or revise conclusions, thereby turning ideas into measurable progress while minimizing hallucinations.
By Yiming Ren, Xiang Liu, Qumeng Sun, Xiao Zhang, Jiahao Li, Haoyang Zhang, Junjie Wang
arXiv:2607. 19044v1 Announce Type: new Abstract: Leveraging large language models (LLMs) for molecular generation has shown remarkable potential in chemical and drug design.
By Mingxuan Ouyang, Hao Lan, Wanyu Lin