arXiv Machine Learning

Research on Domain Information Mining and Theme Evolution of Scientific Papers

arXiv:2204. 08476v2 Announce Type: replace-cross Abstract: In recent years, with the increase of social investment in scientific research, the number of research results in various fields has increased significantly.

arXiv AI
Jul 14

Research on Intellectual Property Resource Profile and Evolution Law

arXiv:2204. 06221v2 Announce Type: replace-cross Abstract: In the era of big data, intellectual property-oriented scientific and technological resources show the trend of large data scale, high information density, and low value density, which brings severe challenges to the effective use of intellectual property resources, and the demand for mining hidden information in intellectual property is increasing.

By Yuhui Wang, Yingxia Shao, Ang Li
arXiv Computation and Language
Sep 2

The Scientific Contribution Graph: Automated Literature-based Technological Roadmapping at Scale

The paper introduces the Scientific Contribution Graph, a large-scale resource that extracts 6 million scientific contributions from 655 k open-access papers across multiple disciplines and links them with 36 million prerequisite edges. It frames automated technological roadmapping as the task of identifying contributions and their prerequisites, and presents a new scientific prerequisite prediction task where models forecast which existing technologies enable future discoveries. The authors report that current models achieve a 0.48 MAP score on temporally-filtered backtesting, indicating rapid progress in this area.

By Peter A. Jansen
arXiv AI
Jun 19

Charting the Future of Scholarly Knowledge with AI: A Community Perspective

arXiv:2509. 02581v2 Announce Type: replace-cross Abstract: Despite the growing availability of tools designed to support scholarly knowledge extraction and organization, many researchers still rely on manual methods, sometimes due to unfamiliarity with existing technologies or limited access to domain-adapted solutions.

By Azanzi Jiomekong, Hande K\"u\c{c}\"uk McGinty, Keith G. Mills, Allard Oelen, Enayat Rajabi, Harry McElroy, Antrea Christou, Anmol Saini, Janice Anta Zebaze, Hannah Kim, Anna M. Jacyszyn, Gollam Rabby, Dirk Betz, Claudia Biniossek, Sanju Tiwari, S\"oren Auer
arXiv Computation and Language
Sep 7

Measuring the Novelty of Biomedical Papers Using the Latent Distances between Knowledge Units

The paper proposes a new method for measuring the novelty of biomedical research by calculating latent distances between knowledge units—specifically MeSH terms—using three types of relationships: network, semantic, and hierarchical. It demonstrates that each relationship captures distinct distances and that combining all three yields a more accurate novelty assessment than existing metrics. Validation on a large PLoS ONE dataset and a H1 Connect dataset shows stronger alignment with peer judgments compared to prior indicators.

By Yi Zhao, Heng Zhang, Yuzhuo Wang, Wenqing Wu, Tong Bao, Chengzhi Zhang