arXiv Computation and Language By Yi Zhao, Heng Zhang, Yuzhuo Wang, Wenqing Wu, Tong Bao, Chengzhi Zhang

Measuring the Novelty of Biomedical Papers Using the Latent Distances between Knowledge Units

Read the original on arXiv Computation and Language →

The paper proposes a new method for measuring the novelty of biomedical research by calculating latent distances between knowledge units—specifically MeSH terms—using three types of relationships: network, semantic, and hierarchical. It demonstrates that each relationship captures distinct distances and that combining all three yields a more accurate novelty assessment than existing metrics. Validation on a large PLoS ONE dataset and a H1 Connect dataset shows stronger alignment with peer judgments compared to prior indicators.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Sep 12

NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment

NovGauge is a new benchmark designed to diagnose large language models’ ability to assess scientific paper novelty. It contains 619 paper pairs and 50 multi-paper sets, each labeled along three dimensions—task, problem, and method—by experts from ICLR reviewer overlap claims and survey co-citations. The study evaluates 18 LLMs, revealing high hallucination rates and weak evidence grounding, with the best model achieving only 43‑72% verified F1 across dimensions.

By Guoqiang Zhang, Kexin Tan, Ming Zhang, Li Ju, Wenqing Jing, Zhonghan Yue, Jiayi Chen, Shiqiang Wu, Shaofan Liu, Yue Zhang, Yuankai Ying, Yang Shi, Tao Gui, Qi Zhang, Xuanjing Huang