The paper introduces LDC, a framework that learns to generate research ideas with dynamic control. It combines supervised fine‑tuning on paper‑idea pairs with controllable reinforcement learning that optimizes novelty, feasibility, and effectiveness. During inference, sentence‑level controllers steer the generation process to balance these dimensions.
By Ruochen Li, Liqiang Jing, Chi Han, Jiawei Zhou, Xinya Du
The paper "Learning to Ideate for Scientific Impact" explores using delayed signals of scientific uptake—specifically citation-normalized impact—as feedback to steer large language models toward generating high‑impact research ideas. The authors build a dataset of over 100,000 computer science papers, train a reward model to predict citation impact from goal‑idea pairs, and align an idea generator via supervised fine‑tuning and reinforcement learning. Evaluation with a reference‑grounded protocol shows that the RL‑tuned model consistently produces ideas with higher estimated impact than baseline models.
By Shubham Kale, Aniketh Garikaparthi, Manasi Patwardhan
arXiv:2606. 09105v1 Announce Type: new Abstract: Generating novel, feasible, and high-quality research ideas is an important yet challenging task in scientific discovery.
By Xu Li, Hanzhe Tu, Xun Han
RATIO (Retrieval Across Typed Ideation Operations) is a large-scale benchmark designed to evaluate how well retrieval systems can support scientific inspiration. It defines relevance through three ideation moves—Address, Broaden, and Specify—each targeting different levels of abstraction in literature retrieval. The benchmark is built from millions of full-text CS papers using a novel discourse-marker distant supervision method, and includes extensive LLM and human vetting to ensure quality.
By Maayan Sharon, Tom Hope
arXiv:2608. 25826v1 Announce Type: cross Abstract: A recent line of synthetic-data work reconstructs the thinking behind existing text rather than rewriting the text itself, but it operates on short web passages, recovers only local thoughts, and leaves the structure of whole documents untouched.
By Qiankai Xu, Qiguang Chen, Zixin Su, Wenhao Huang, Yue Gao, Jiaheng Liu, Ge Zhang
A recent line of synthetic-data work reconstructs the thinking behind existing text rather than rewriting the text itself, but it operates on short web passages, recovers only local thoughts, and leav...
Highlights provide a concise summary of the main contributions of an academic paper and help readers quickly understand its focus. However, many journals do not provide highlights, which limits their use in literature retrieval, text mining, and bibliometric analysis.
ConvergeWriter introduces a bottom‑up, data‑driven framework for long‑form document generation that first retrieves exhaustive knowledge from a source corpus and clusters it into distinct knowledge groups. These clusters then guide the creation of a hierarchical outline and the final text, ensuring the output is strictly grounded in the retrieved material and traceable to its sources. Experiments on 14B and 32B LLMs show that this approach matches or surpasses state‑of‑the‑art baselines, especially in scenarios requiring high factual fidelity and structural coherence.
By Binquan Ji, Jiaqi Wang, Ruiting Li, Xingchen Han, Yiyang Qi, Shichao Wang, Yifei Lu, Yuantao Han, Feiliang Ren
arXiv:2508.14273v3 Announce Type: replace
Abstract: As researchers increasingly adopt LLMs as writing assistants, generating high-quality research paper introductions remains both challenging and ess...
By Krishna Garg, Firoz Shaik, Sambaran Bandyopadhyay, Cornelia Caragea
arXiv:2609.22104v1 Announce Type: new
Abstract: As automated scientific discovery advances, Large Language Models (LLMs) can now generate research ideas at an unprecedented scale, shifting the bottle...
By Rongcan Pei, Fang Guo, Qinglin Qi, Qi Zhu, Yun Luo, Jianhao Yan, Minjun Zhu, Qiujie Xie, Dehong Zheng, Yue Zhang
LitPivot is a tool that supports researchers in developing well-situated research ideas by dynamically linking literature exploration with idea refinement. It introduces literature-initiated pivots, where engaging with relevant papers prompts revisions to an idea, and vice versa, updating the set of pertinent literature. In a lab study with 17 participants, users produced higher-rated ideas and reported a stronger understanding of the literature space, while an open-ended study with five participants illustrated how LitPivot facilitates iterative idea evolution.
By Hita Kambhamettu, Bhavana Dalvi Mishra, Andrew Head, Jonathan Bragg, Aakanksha Naik, Joseph Chee Chang, Pao Siangliulue
arXiv:2609.37891v1 Announce Type: cross
Abstract: Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--...
By Pierre-Carl Langlais, Pieter Delobelle, Yannick Detrois, Pavel Chizhov, Carlos Rosas-Hinostroza, Neil Si Smail, Benjamin Burtin, Hanna Shcharbakova, Ivan Yamshchikov, Anastasia Stasenko