NovGauge is a new benchmark designed to diagnose large language models’ ability to assess scientific paper novelty. It contains 619 paper pairs and 50 multi-paper sets, each labeled along three dimensions—task, problem, and method—by experts from ICLR reviewer overlap claims and survey co-citations. The study evaluates 18 LLMs, revealing high hallucination rates and weak evidence grounding, with the best model achieving only 43‑72% verified F1 across dimensions.
By Guoqiang Zhang, Kexin Tan, Ming Zhang, Li Ju, Wenqing Jing, Zhonghan Yue, Jiayi Chen, Shiqiang Wu, Shaofan Liu, Yue Zhang, Yuankai Ying, Yang Shi, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv:2606. 12071v1 Announce Type: cross Abstract: LLMs are increasingly used to generate and judge scientific ideas.
By Soumitra Sinhahajari, Navonil Majumder, Soujanya Poria
arXiv:2609.23264v1 Announce Type: new
Abstract: Peer-review evaluation is increasingly being automated with LLM-as-a-judge metrics, but this creates a measurement risk. A review may receive a high sc...
By Shakiba Amirshahi, Sajad Ebrahimi, Hai Son Le, Negar Arabzadeh, Ebrahim Bagheri
arXiv:2609.25046v1 Announce Type: new
Abstract: Peer review plays a central role in scholarly publishing, yet verifying whether reviewer claims are supported by manuscript evidence remains a largely...
By Alireza Daghighfarsoodeh, Sajad Ebrahimi, Ali Ghorbanpour, Soroush Sadeghian, Radin Cheraghi, Negar Arabzadeh, Ebrahim Bagheri
PaperDoctor is an agent framework that provides evidence‑grounded, actionable feedback for scientific papers before submission. It evaluates writing, layout, references, code, theory, prior work, and experiments through a three‑layer hierarchical system, linking each critique to specific evidence and revision suggestions. The system selectively rebuilds and reruns experiments to uncover reproducibility gaps, and an interactive interface lets authors explore findings tied to their manuscript.
By Kevin Qinghong Lin, Siyuan Hu, Pan Lu, Yu Chen, Yanzhe Chen, Owen Queen, Yupeng Chen, Jialin Yu, Junchi Yu, Zifeng Ding, Yuanfeng Ji, Sheng Liu, Jindong Gu, Linjie Li, Mike Zheng Shou, Philip Torr, James Zou
The paper investigates language-of-study (LoS) bias in NLP peer reviews, defining and distinguishing negative and positive forms of bias. Using a new dataset, LOBSTER, and an LLM-based detection pipeline, the authors analyze 15,645 reviews and find that non‑English papers experience significantly higher bias rates, with negative bias outweighing positive bias. They further identify four subcategories of negative bias, noting that demanding unjustified cross‑lingual generalization is the most common.
By Ehsan Barkhordar, Abdulfattah Safa, Verena Blaschke, Erika Lombart, Marie-Catherine de Marneffe, G\"ozde G\"ul \c{S}ahin