arXiv AI

Measuring the Gap Between Human and LLM Research Ideas

arXiv:2607. 01233v1 Announce Type: cross Abstract: LLMs are increasingly used to brainstorm research ideas, but existing evaluations mostly judge individual ideas by novelty, feasibility, or expert preference.

arXiv AI
Sep 24

LitPivot: Developing Well-Situated Research Ideas Through Dynamic Contextualization and Critique within the Literature Landscape

LitPivot is a tool that supports researchers in developing well-situated research ideas by dynamically linking literature exploration with idea refinement. It introduces literature-initiated pivots, where engaging with relevant papers prompts revisions to an idea, and vice versa, updating the set of pertinent literature. In a lab study with 17 participants, users produced higher-rated ideas and reported a stronger understanding of the literature space, while an open-ended study with five participants illustrated how LitPivot facilitates iterative idea evolution.

By Hita Kambhamettu, Bhavana Dalvi Mishra, Andrew Head, Jonathan Bragg, Aakanksha Naik, Joseph Chee Chang, Pao Siangliulue
arXiv AI
Sep 25

Learning to Ideate for Scientific Impact

The paper "Learning to Ideate for Scientific Impact" explores using delayed signals of scientific uptake—specifically citation-normalized impact—as feedback to steer large language models toward generating high‑impact research ideas. The authors build a dataset of over 100,000 computer science papers, train a reward model to predict citation impact from goal‑idea pairs, and align an idea generator via supervised fine‑tuning and reinforcement learning. Evaluation with a reference‑grounded protocol shows that the RL‑tuned model consistently produces ideas with higher estimated impact than baseline models.

By Shubham Kale, Aniketh Garikaparthi, Manasi Patwardhan
arXiv AI
Sep 25

Generating Interesting Scientific Ideas using Knowledge Graphs and LLMs: Evaluations with 100 Research Group Leaders

The paper introduces SciMuse, an AI system that generates personalized research ideas by combining a knowledge graph of 58 million papers with a large language model. A large-scale evaluation involving over 100 research group leaders across disciplines rated more than 4,400 ideas, yielding modest overall interest scores but showing that 24.9% were rated highly. The study also demonstrates that graph-derived features can predict idea interest and can be used to control idea properties, offering a new methodology for generating and assessing scientific ideas.

By Xuemei Gu, Mario Krenn
arXiv Computation and Language
Aug 28

RATIO: A Benchmark for Retrieval Across Typed Ideation Operations in Scientific Literature

RATIO (Retrieval Across Typed Ideation Operations) is a large-scale benchmark designed to evaluate how well retrieval systems can support scientific inspiration. It defines relevance through three ideation moves—Address, Broaden, and Specify—each targeting different levels of abstraction in literature retrieval. The benchmark is built from millions of full-text CS papers using a novel discourse-marker distant supervision method, and includes extensive LLM and human vetting to ensure quality.

By Maayan Sharon, Tom Hope