arXiv Computation and Language

Measuring the Novelty of Biomedical Papers Using the Latent Distances between Knowledge Units

The paper proposes a new method for measuring the novelty of biomedical research by calculating latent distances between knowledge units—specifically MeSH terms—using three types of relationships: network, semantic, and hierarchical. It demonstrates that each relationship captures distinct distances and that combining all three yields a more accurate novelty assessment than existing metrics. Validation on a large PLoS ONE dataset and a H1 Connect dataset shows stronger alignment with peer judgments compared to prior indicators.

arXiv AI
Sep 12

NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment

NovGauge is a new benchmark designed to diagnose large language models’ ability to assess scientific paper novelty. It contains 619 paper pairs and 50 multi-paper sets, each labeled along three dimensions—task, problem, and method—by experts from ICLR reviewer overlap claims and survey co-citations. The study evaluates 18 LLMs, revealing high hallucination rates and weak evidence grounding, with the best model achieving only 43‑72% verified F1 across dimensions.

By Guoqiang Zhang, Kexin Tan, Ming Zhang, Li Ju, Wenqing Jing, Zhonghan Yue, Jiayi Chen, Shiqiang Wu, Shaofan Liu, Yue Zhang, Yuankai Ying, Yang Shi, Tao Gui, Qi Zhang, Xuanjing Huang
arXiv AI
Sep 7

Hakken: Predicting future discoveries to fill the gaps in today's knowledge

Hakken is a domain‑agnostic system that predicts and explains future scientific discoveries by combining transformer‑based models trained on temporal knowledge graphs with large language model semantic knowledge. It identifies novel relationships between scientific concepts that extend beyond the deductive hull of existing knowledge and provides explanations to help scientists assess these predictions. In the biomedical domain, Hakken set a new benchmark for time‑aware multi‑label relation prediction, generated 1.5 million high‑confidence hypotheses about aging, and experimentally confirmed two predictions that revealed previously undocumented interactions relevant to drug discovery.

By Tarek R. Besold, Uchenna Akujuobi, Pablo Sanchez, Alessandra Toniato, Kana Maruyama, Jihun Choi, Samy Badreddine, Frederick Gifford, Daniel Evans-Yamamoto, Sucheendra K. Palaniappan, Miquel Ferrer, Kae Nagano, Iris Rossell, Tom Joy, Hatem ElShazly, Chrysa Iliopoulou, Christoph Wehner, Thiviyan Thanapalasingam, Susana Nunes, Pedro G. Cotovio, Peter Wurman, Peter Stone, Hiroaki Kitano, Michael Spranger
arXiv Computation and Language
Aug 27

Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty

The paper introduces Think‑Probe‑Respond (TPR), a lightweight method to improve large language models’ ability to judge the novelty of research ideas. It identifies a systematic bias where models tend to label ideas as "medium novel" despite generating human‑like rationales, and shows that probing hidden states during reasoning and conditioning the final response on these probes boosts novelty judgment accuracy by 22.30%. TPR effectively reduces the medium‑novelty bias across strong baseline models.

By Tim Schopf, Tobias Schreieder, Akiko Aizawa