arXiv AI

PreScience: A Dataset and Benchmark for Scientific Forecasting

arXiv:2602. 20459v2 Announce Type: replace Abstract: Can AI systems trained on the existing scientific record forecast the advances that will follow?

arXiv AI
Sep 25

Learning to Ideate for Scientific Impact

The paper "Learning to Ideate for Scientific Impact" explores using delayed signals of scientific uptake—specifically citation-normalized impact—as feedback to steer large language models toward generating high‑impact research ideas. The authors build a dataset of over 100,000 computer science papers, train a reward model to predict citation impact from goal‑idea pairs, and align an idea generator via supervised fine‑tuning and reinforcement learning. Evaluation with a reference‑grounded protocol shows that the RL‑tuned model consistently produces ideas with higher estimated impact than baseline models.

By Shubham Kale, Aniketh Garikaparthi, Manasi Patwardhan
arXiv Computation and Language
Sep 2

The Scientific Contribution Graph: Automated Literature-based Technological Roadmapping at Scale

The paper introduces the Scientific Contribution Graph, a large-scale resource that extracts 6 million scientific contributions from 655 k open-access papers across multiple disciplines and links them with 36 million prerequisite edges. It frames automated technological roadmapping as the task of identifying contributions and their prerequisites, and presents a new scientific prerequisite prediction task where models forecast which existing technologies enable future discoveries. The authors report that current models achieve a 0.48 MAP score on temporally-filtered backtesting, indicating rapid progress in this area.

By Peter A. Jansen
arXiv AI
2d ago

ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

ScholarCatalyst is a new benchmark that evaluates how well AI systems can retrieve research papers that inspire new work. The dataset was created by having 184 lead authors of 207 recent computer science papers annotate which earlier papers helped their projects, providing detailed rationales. The benchmark tests retrieval from the literature available at the start of a project, revealing that current agentic search and even advanced models like Claude Fable 5.1 perform only modestly better than simple embedding retrieval.

By Sohyeon Kim, Yoonho Lee, Bo Liu, Dayoon Ko, Rulin Shao, Seungone Kim, Graham Neubig, Pang Wei Koh, Aakanksha Chowdhery, Akari Asai, Omar Khattab, Yejin Choi, Gunhee Kim, Chelsea Finn