arXiv AI

Towards Nexus-Score: Metadata Gaps Limit Scholarly AI Attribution

arXiv:2607. 22684v1 Announce Type: cross Abstract: Artificial intelligence systems increasingly mediate how science is found and credited.

arXiv AI
Sep 23

ScholarStack: Layered Research Asset Orchestration and Cross-Task Reuse for Scientific Agents

arXiv:2609.23735v2 Announce Type: new Abstract: Scientific agents support a range of literature-based research tasks, such as retrieval, question answering, evidence-grounded generation, and claim as...

By ScholarSeed AI Team, Caoqinwei Gong, Xue Jiang, Wei Luo, Xiaoyu Qiu, Jiayi Sheng, Yi Wang, Zheng Yu, Ao Zhang, Haifan Zhang, Hanwei Zhang, Jihai Zhang, Yuan Cao, Wei Chen, Liyun Dai, Wenkai Fang, Guanglei Wang, Kai Ying, Tingyu Zhu, Wotao Yin
arXiv Machine Learning
Jul 30

Can AI agents conduct open-ended AI research? Early evidence from two case studies

arXiv:2607. 27191v1 Announce Type: cross Abstract: Forecasts of explosive AI progress hinge on AI agents automating AI research.

By Peter Kirgis, Sayash Kapoor, Andrew Schwartz, Stephan Rabanser, David Africa, Konstantinos Voudouris, Viet Nguyen, Toby Pilditch, Magda Dubois, Harry Coppock, Cozmin Ududec, Nitya Nadgir, Matilda Orona, Tilman Bayer, Derrick Chan-Sew, Yue Ling, Abhishek Shetty, Helen Toner, Gillian Hadfield, Seth Lazar, Steve Newman, Shoshannah Tekofsky, Rishi Bommasani, Arvind Narayanan
arXiv AI
Aug 28

6.5% of the Neuro-Symbolic Literature Can Be Reproduced from Its Published Artifacts, a Six-Stage Audit Framework and First Instantiation

The paper introduces a six‑stage audit framework for assessing reproducibility in computer science literature and applies it to the neuro‑symbolic AI (NSAI) subfield. Using the framework, the authors screened 5,497 records, identified 1,304 eligible studies, and found verifiable code artifacts for only 455 of them. Of those, they fully or partially reproduced 85 studies, representing 6.52% of the eligible corpus and 18.68% of attempted reruns, highlighting a significant reproducibility gap even when code is declared available.

By Brandon Colelough, Vladimir Martirosyan, Ishan Tamrakar, William Regli, Aditya Kumar, Anh N. Nhu, Dhruv Dubey, Raj Ambavane, Haowei Deng
arXiv AI
Sep 10

Attribution in Scientific Literature: New Benchmark and Methods

The paper introduces REASONS, a benchmark of 12,723 sentence-level citation instances across 12 arXiv subject categories, to evaluate scientific citation attribution under different evidence conditions. It proposes a dual-metric framework—Abstention Rate (AR) and Hallucination Rate (HR)—to balance reliability and responsiveness. Experiments with proprietary and open-source LLMs across various prompting and retrieval settings show that advanced Retrieval-Augmented Generation (RAG) reduces hallucinations but increases abstention, while adversarial metadata can push hallucination rates above 85%. Human evaluation confirms a high ratio of factual hallucinations to acceptable paraphrases, underscoring the need for systems that can appropriately abstain under uncertainty.

By Deepa Tilwani, Yash Saxena, Seyedali Mohammadi, Ankur Padia, Edward Raff, Amit Sheth, Srinivasan Parthasarathy, Manas Gaur