arXiv AI

ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research

ScholarCatalyst is a new benchmark that evaluates how well AI systems can retrieve research papers that inspire new work. The dataset was created by having 184 lead authors of 207 recent computer science papers annotate which earlier papers helped their projects, providing detailed rationales. The benchmark tests retrieval from the literature available at the start of a project, revealing that current agentic search and even advanced models like Claude Fable 5.1 perform only modestly better than simple embedding retrieval.

arXiv Computation and Language
Aug 28

RATIO: A Benchmark for Retrieval Across Typed Ideation Operations in Scientific Literature

RATIO (Retrieval Across Typed Ideation Operations) is a large-scale benchmark designed to evaluate how well retrieval systems can support scientific inspiration. It defines relevance through three ideation moves—Address, Broaden, and Specify—each targeting different levels of abstraction in literature retrieval. The benchmark is built from millions of full-text CS papers using a novel discourse-marker distant supervision method, and includes extensive LLM and human vetting to ensure quality.

By Maayan Sharon, Tom Hope
arXiv Computation and Language
Sep 16

PaperDoctor: Evidence-Grounded and Actionable Feedback for Scientific Papers in Progress

PaperDoctor is an agent framework that provides evidence‑grounded, actionable feedback for scientific papers before submission. It evaluates writing, layout, references, code, theory, prior work, and experiments through a three‑layer hierarchical system, linking each critique to specific evidence and revision suggestions. The system selectively rebuilds and reruns experiments to uncover reproducibility gaps, and an interactive interface lets authors explore findings tied to their manuscript.

By Kevin Qinghong Lin, Siyuan Hu, Pan Lu, Yu Chen, Yanzhe Chen, Owen Queen, Yupeng Chen, Jialin Yu, Junchi Yu, Zifeng Ding, Yuanfeng Ji, Sheng Liu, Jindong Gu, Linjie Li, Mike Zheng Shou, Philip Torr, James Zou
arXiv AI
Aug 19

The Problem Is the Problem: Towards Scalable Mathematical Discovery

The paper introduces a new human‑AI collaboration paradigm for mathematical discovery, shifting from selecting individual problems to exploring broad research directions. It presents the Find, Attempt, and Recommend (FAR) pipeline, which automatically searches a literature corpus, filters candidate conjectures, and surfaces promising resolutions for expert review. In a combinatorics pilot, FAR processed over 5,000 papers, identified thousands of open conjectures, and ultimately highlighted 77 items that led to new discoveries.

By Zeyu Zheng, Shengtong Zhang, Jeremy Avigad, Prasad Tetali, Sean Welleck
arXiv AI
3d ago

Large Knowledge Model: A Knowledge Foundation for Agentic Science at Scale

arXiv:2609.27297v2 Announce Type: replace Abstract: Agentic science envisions many autonomous agents investigating concurrently while building on a shared, evolving body of scientific knowledge. This...

By Yuan Huang, Sihan Hu, Hongyu Gu, Chao Ma, Jiaxing Zhang, Zhiyong Zou, Caiyu Fan, Yan Xiao, Mingjun Xu, Chenyu Xie, Mingzhen Ju, Zhehao Ma, Qi Zhang, Baozong Wang, Yu Li, Zhiyuan Yao, Ruoxue Liao, Xinyu Li, Linfeng Zhang, Kun Chen, Weinan E