NovGauge is a new benchmark designed to diagnose large language models’ ability to assess scientific paper novelty. It contains 619 paper pairs and 50 multi-paper sets, each labeled along three dimensions—task, problem, and method—by experts from ICLR reviewer overlap claims and survey co-citations. The study evaluates 18 LLMs, revealing high hallucination rates and weak evidence grounding, with the best model achieving only 43‑72% verified F1 across dimensions.
By Guoqiang Zhang, Kexin Tan, Ming Zhang, Li Ju, Wenqing Jing, Zhonghan Yue, Jiayi Chen, Shiqiang Wu, Shaofan Liu, Yue Zhang, Yuankai Ying, Yang Shi, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv:2604. 15145v2 Announce Type: replace Abstract: The rigorous evaluation of the novelty of a scientific paper is, even for human scientists, a challenging task.
By Miri Liu, ChengXiang Zhai
arXiv:2609.22104v1 Announce Type: new
Abstract: As automated scientific discovery advances, Large Language Models (LLMs) can now generate research ideas at an unprecedented scale, shifting the bottle...
By Rongcan Pei, Fang Guo, Qinglin Qi, Qi Zhu, Yun Luo, Jianhao Yan, Minjun Zhu, Qiujie Xie, Dehong Zheng, Yue Zhang
arXiv:2605.29522v2 Announce Type: replace
Abstract: As scientific literature grows rapidly and research increasingly involves AI agents, automated survey generation has become a key capability for bo...
By Ziyue Yang, Da Ma, Hanqi Li, Zijian Wang, Tiancheng Huang, Zijian Hu, Chenrun Wang, Yunzhe Zhang, Xiaobao Wu, Kai Yu, Lu Chen
arXiv:2607. 05456v1 Announce Type: new Abstract: While recent advances in large language models have enabled end-to-end automated manuscript generation, existing systems suffer from three critical deficiencies: (i) generated claims are not deterministically grounded in verifiable literature, (ii) experimental results are frequently fabricated rather than executed, and (iii) there exists no standardized, multi-dimensional framework to assess whether AI-generated manuscripts meet the quality and rigor required for real-world publication.
By Ramsha Kamran, Maheera Amjad, Zartasha Mustansar, Arsalan Shaukat, Salma Sherbaz, Muhammad U. S. Khan
PaperDoctor is an agent framework that provides evidence‑grounded, actionable feedback for scientific papers before submission. It evaluates writing, layout, references, code, theory, prior work, and experiments through a three‑layer hierarchical system, linking each critique to specific evidence and revision suggestions. The system selectively rebuilds and reruns experiments to uncover reproducibility gaps, and an interactive interface lets authors explore findings tied to their manuscript.
By Kevin Qinghong Lin, Siyuan Hu, Pan Lu, Yu Chen, Yanzhe Chen, Owen Queen, Yupeng Chen, Jialin Yu, Junchi Yu, Zifeng Ding, Yuanfeng Ji, Sheng Liu, Jindong Gu, Linjie Li, Mike Zheng Shou, Philip Torr, James Zou
arXiv:2606. 28277v1 Announce Type: cross Abstract: Artificial intelligence is driving a revolution in scientific discovery, accelerating everything from hypothesis generation to mathematical theorem proving.
By Rajesh Jayaram, Drew Tyler, David Woodruff, Corinna Cortes, Yossi Matias, Vahab Mirrokni, Vincent Cohen-Addad
arXiv:2609.13760v1 Announce Type: new
Abstract: Publishing a research manuscript is a routine yet demanding part of scientific life: time-consuming, stressful, and often uncertain in outcome. Recent...
By Jiawen Chen, Zichen Zhang, Bingxuan Li, Quan Sun, Yiyan Zhang, Edric Tam, Jinjie Lin, Didong Li, Yun Li, Bingxin Zhao
arXiv:2606. 25198v2 Announce Type: replace Abstract: Autonomous AI Research promises to accelerate the scientific progress of machine learning.
By Antonis Antoniades, Deepak Nathani, Ritam Saha, Alfonso Amayuelas, Ivan Bercovich, Zhaotian Weng, Vignesh Baskaran, Kunal Bhatia, William Yang Wang
PaperScout is an autonomous agent that treats academic paper search as a sequential decision-making process, dynamically deciding when and how to use search and expansion tools based on accumulated context. The authors identify a granularity mismatch in standard reinforcement learning for multi-turn tasks and propose Proximal Sequence Policy Optimization (PSPO), a sequence-level policy optimization method that aligns learning with agent–environment interactions. Experiments on synthetic and real-world benchmarks show that PaperScout outperforms structured retrieval and RL baselines in recall and relevance, demonstrating the effectiveness of its adaptive agentic framework and optimization strategy.
By Tingyue Pan, Jie Ouyang, Mingyue Cheng, Qingchuan Li, Zirui Liu, Daoyu Wang, Mingfan Pan, Shuo Yu, Qi Liu, Enhong Chen
Nomad is an autonomous system designed to explore and discover insights within large data corpora. It builds an explicit Exploration Map to systematically traverse a domain, generating and testing hypotheses with an explorer agent that leverages document, web, and database searches. After verification, it produces cited reports and meta-reports, and its evaluation framework assesses trustworthiness, quality, and diversity, showing superior performance over baselines on UN, WHO, and arXiv datasets.
By Bokang Jia, Samta Kamboj, Satheesh Katipomu, Seung Hun Han, Neha Sengupta, Andrew Jackson
The paper introduces BioCheck Agent, an LLM-based system that generates structured biomedical fact‑checking reports using agentic search and a reinforcement‑learning framework called EG‑GRPO. Unlike prior methods that output only supported or refuted labels, BioCheck Agent synthesizes conclusions with retrieved evidence from PubMed, employing advanced Boolean search operators. Experiments show that, compared to the base Qwen3.5‑4B model, BioCheck Agent improves label prediction accuracy on SciFact by 9.95 %, raises evidence quality by 3.7 %, and reduces hallucinations by 19.63 %.
By Jiongxiao Wang, Dingli Ma, Chaoqun Ni