arXiv AI

RWGBench: Evaluating Scholarly Positioning in Related Work Generation

arXiv:2606. 24894v2 Announce Type: replace-cross Abstract: Large language models have shown strong fluency in scientific writing, yet the evaluation of related work generation (RWG) remains limited.

arXiv AI
Jun 9

GIScholarBench: Benchmarking LLM Overconfidence in GIS Research

arXiv:2606. 08036v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used in academic research workflows, but scholarly tasks require high factual precision and therefore expose a key weakness: overconfidence.

By Zongrng Li, Mingzheng Yang, Lei Zou, Hongxu Ma, Hao Tian, Siqi Zhou, Wenjing Gong, Kaili Zhang, Bingqian Chen, Mitch Zhang, Yifan Yang
arXiv AI
3d ago

Re-ranking and Late Interaction Drive Retrieval Quality: A Controlled Comparison of RAG Strategies for Scientific Question Answering

The paper presents a controlled comparison of six retrieval-augmented generation (RAG) strategies for scientific question answering on a large arXiv corpus. All pipelines use the same LLM generator and evaluation protocol, differing only in retrieval design—ranging from classic dense retrieval to late‑interaction methods like ColBERTv2. The authors also release a synthetic question dataset and code to enable reproducible, large‑scale evaluation of RAG trade‑offs.

By Bhagyesh Rathi, Eshan Chawla, William B. Andreopoulos
arXiv AI
Sep 25

Learning to Ideate for Scientific Impact

The paper "Learning to Ideate for Scientific Impact" explores using delayed signals of scientific uptake—specifically citation-normalized impact—as feedback to steer large language models toward generating high‑impact research ideas. The authors build a dataset of over 100,000 computer science papers, train a reward model to predict citation impact from goal‑idea pairs, and align an idea generator via supervised fine‑tuning and reinforcement learning. Evaluation with a reference‑grounded protocol shows that the RL‑tuned model consistently produces ideas with higher estimated impact than baseline models.

By Shubham Kale, Aniketh Garikaparthi, Manasi Patwardhan
arXiv AI
Jul 28

VecTree-RAG: An Agentic Retrieval-Augmented Generation Framework Combining Vector and Tree Retrieval for Efficiency and Accuracy

arXiv:2607. 23006v1 Announce Type: cross Abstract: Scientific question answering requires a retrieval system to solve two distinct problems: identifying which papers are relevant and locating the supporting evidence within those papers.

By Xinyan Zhong, Yuwei Shi, Yuqi Wei, Chen Shen, Tianhang Zhou, Zhenghao Wu
arXiv AI
Sep 23

ScholarStack: Layered Research Asset Orchestration and Cross-Task Reuse for Scientific Agents

arXiv:2609.23735v2 Announce Type: new Abstract: Scientific agents support a range of literature-based research tasks, such as retrieval, question answering, evidence-grounded generation, and claim as...

By ScholarSeed AI Team, Caoqinwei Gong, Xue Jiang, Wei Luo, Xiaoyu Qiu, Jiayi Sheng, Yi Wang, Zheng Yu, Ao Zhang, Haifan Zhang, Hanwei Zhang, Jihai Zhang, Yuan Cao, Wei Chen, Liyun Dai, Wenkai Fang, Guanglei Wang, Kai Ying, Tingyu Zhu, Wotao Yin
arXiv Computation and Language
Aug 28

RATIO: A Benchmark for Retrieval Across Typed Ideation Operations in Scientific Literature

RATIO (Retrieval Across Typed Ideation Operations) is a large-scale benchmark designed to evaluate how well retrieval systems can support scientific inspiration. It defines relevance through three ideation moves—Address, Broaden, and Specify—each targeting different levels of abstraction in literature retrieval. The benchmark is built from millions of full-text CS papers using a novel discourse-marker distant supervision method, and includes extensive LLM and human vetting to ensure quality.

By Maayan Sharon, Tom Hope