arXiv AI By Miri Liu, ChengXiang Zhai

An Axiomatic Benchmark for Evaluation of Scientific Novelty Metrics

Read the original on arXiv AI →

arXiv:2604. 15145v2 Announce Type: replace Abstract: The rigorous evaluation of the novelty of a scientific paper is, even for human scientists, a challenging task.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 12

NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment

NovGauge is a new benchmark designed to diagnose large language models’ ability to assess scientific paper novelty. It contains 619 paper pairs and 50 multi-paper sets, each labeled along three dimensions—task, problem, and method—by experts from ICLR reviewer overlap claims and survey co-citations. The study evaluates 18 LLMs, revealing high hallucination rates and weak evidence grounding, with the best model achieving only 43‑72% verified F1 across dimensions.

By Guoqiang Zhang, Kexin Tan, Ming Zhang, Li Ju, Wenqing Jing, Zhonghan Yue, Jiayi Chen, Shiqiang Wu, Shaofan Liu, Yue Zhang, Yuankai Ying, Yang Shi, Tao Gui, Qi Zhang, Xuanjing Huang