arXiv AI

Learning to Ideate for Scientific Impact

The paper "Learning to Ideate for Scientific Impact" explores using delayed signals of scientific uptake—specifically citation-normalized impact—as feedback to steer large language models toward generating high‑impact research ideas. The authors build a dataset of over 100,000 computer science papers, train a reward model to predict citation impact from goal‑idea pairs, and align an idea generator via supervised fine‑tuning and reinforcement learning. Evaluation with a reference‑grounded protocol shows that the RL‑tuned model consistently produces ideas with higher estimated impact than baseline models.

arXiv AI
Sep 4

LDC: Learning to Generate Research Idea with Dynamic Control

The paper introduces LDC, a framework that learns to generate research ideas with dynamic control. It combines supervised fine‑tuning on paper‑idea pairs with controllable reinforcement learning that optimizes novelty, feasibility, and effectiveness. During inference, sentence‑level controllers steer the generation process to balance these dimensions.

By Ruochen Li, Liqiang Jing, Chi Han, Jiawei Zhou, Xinya Du
arXiv Computation and Language
Aug 28

RATIO: A Benchmark for Retrieval Across Typed Ideation Operations in Scientific Literature

RATIO (Retrieval Across Typed Ideation Operations) is a large-scale benchmark designed to evaluate how well retrieval systems can support scientific inspiration. It defines relevance through three ideation moves—Address, Broaden, and Specify—each targeting different levels of abstraction in literature retrieval. The benchmark is built from millions of full-text CS papers using a novel discourse-marker distant supervision method, and includes extensive LLM and human vetting to ensure quality.

By Maayan Sharon, Tom Hope
arXiv AI
5d ago

Generating Interesting Scientific Ideas using Knowledge Graphs and LLMs: Evaluations with 100 Research Group Leaders

The paper introduces SciMuse, an AI system that generates personalized research ideas by combining a knowledge graph of 58 million papers with a large language model. A large-scale evaluation involving over 100 research group leaders across disciplines rated more than 4,400 ideas, yielding modest overall interest scores but showing that 24.9% were rated highly. The study also demonstrates that graph-derived features can predict idea interest and can be used to control idea properties, offering a new methodology for generating and assessing scientific ideas.

By Xuemei Gu, Mario Krenn
arXiv AI
Sep 10

From Citations to Contributions: LLM-Assisted Credit Scoring of Research Articles

The paper proposes a new method for credit scoring research articles that distinguishes between a paper’s original contribution and the prior work it builds upon. It introduces a hierarchical ‘contribution tree’ framework that conserves importance across a document’s structure and separates original from citation-derived credit. Large language models are employed as noisy comparative estimators to scale the analysis, and the approach is extended to collections of articles via weighted citation graphs to produce corpus-level contributions and normalized influence scores.

By Sana Ebrahimi, Suraj Shetiya, Abolfazl Asudeh