arXiv:2606. 12071v1 Announce Type: cross Abstract: LLMs are increasingly used to generate and judge scientific ideas.
By Soumitra Sinhahajari, Navonil Majumder, Soujanya Poria
NovGauge is a new benchmark designed to diagnose large language models’ ability to assess scientific paper novelty. It contains 619 paper pairs and 50 multi-paper sets, each labeled along three dimensions—task, problem, and method—by experts from ICLR reviewer overlap claims and survey co-citations. The study evaluates 18 LLMs, revealing high hallucination rates and weak evidence grounding, with the best model achieving only 43‑72% verified F1 across dimensions.
By Guoqiang Zhang, Kexin Tan, Ming Zhang, Li Ju, Wenqing Jing, Zhonghan Yue, Jiayi Chen, Shiqiang Wu, Shaofan Liu, Yue Zhang, Yuankai Ying, Yang Shi, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv:2604. 15145v2 Announce Type: replace Abstract: The rigorous evaluation of the novelty of a scientific paper is, even for human scientists, a challenging task.
By Miri Liu, ChengXiang Zhai
arXiv:2607. 26066v1 Announce Type: cross Abstract: The growing volume of scientific submissions has motivated interest in using large language models (LLMs) to assist peer review.
By Ranjitha Shivaprasad Ballakuraya, Arash Mahyari, Ashok Srinivasan
arXiv:2606. 08251v1 Announce Type: cross Abstract: Bold projections that artificial intelligence will accelerate scientific discovery have raced ahead of evidence from working scientists, and the field still lacks large-scale, scientist-in-the-loop tests of these claims.
By Honglin Bao, Siyang Wu, Xiao Liu, Sida Li, Shiyun Cao, James A. Evans
arXiv:2609.14057v1 Announce Type: new
Abstract: Frontier LLMs are increasingly capable of conducting automated research, yet their creativity in this setting has not been systematically evaluated. In...
By Yiheng Zhao, Mengzhuo Chen, Chengming Hu, Pengyi Liao, Yiran Pang