SciNLP is a new benchmark dataset for full‑text entity and relation extraction in the NLP domain, comprising 60 manually annotated papers with 6,429 entities and 1,649 relations. It is the first dataset to provide full‑text annotations of entities and their relationships specifically for NLP literature. Experiments show that models trained on SciNLP outperform baselines on certain tasks, and the dataset enabled the automatic construction of a fine‑grained knowledge graph with an average node degree of 3.3.
By Decheng Duan, Yingyi Zhang, Jitong Peng, Chengzhi Zhang
The paper introduces a multilingual, multi-functional framework for disambiguating funder names in scientific publications, using a training dataset that merges the Research Organization Registry with Web of Science and Crossref Open Funder Registry data. By applying multi-task learning with contrastive and multiple negatives ranking losses, the authors fine‑tune open‑weight embedding models from the Sentence Transformer, Gemma, and Qwen3 families, achieving over 90% accuracy in matching Web of Science funder names to ROR identifiers and surpassing general‑purpose LLMs by more than 0.1. For funders not present in ROR, a similarity network is constructed to identify clusters, and the study discusses challenges related to smaller and non‑English‑speaking funders.
By Kanyao Han, Zhiwen You, Jinseok Kim, Jana Diesner
arXiv:2608. 19201v1 Announce Type: cross Abstract: Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific literature are often inconsistent and difficult to systematically identify at scale.
By Hao Xuan, Rithvij Pasupuleti, Ben Liu, Haishuo Sun, Jun Zhang, Zijun Yao, Cuncong Zhong
arXiv:2608. 07254v1 Announce Type: cross Abstract: The increasing specialization of scientific research challenges existing classification systems, which provide effective representations of broad disciplines and research topics but often fail to capture the fine-grained conceptual structure of contemporary science.
By Daniele Raimondi, Feichi Lu, Oliver Grun, Mariia Eremina, Andrea Perlato
The paper introduces a weakly supervised framework for extracting dataset mentions from forced displacement and Fragile, Conflict, and Violence (FCV) documents. It uses a lightweight model trained on general research literature to generate candidate mentions, which are then refined by a large language model that validates or rejects them and corrects boundaries. The refined annotations are augmented with synthetic and contrastive examples to fine‑tune the model, achieving 74.1% precision and 70.5% recall on a benchmark of 1,706 passages, with higher precision (89.5%) on passages that contain dataset references.
By Rafael Macalaba, Aivin V. Solatorio, Patrick Michael Brock, Olivier Dupriez
arXiv:2510. 16152v2 Announce Type: replace-cross Abstract: Scientific literature is increasingly fragmented by disciplinary boundaries, specialized terminology, and potentially sparse keyword systems, making it difficult to capture the evolving structure of modern science.
By Mason Smetana, Lev Khazanovich
arXiv:2606. 07611v1 Announce Type: cross Abstract: This paper proposes an improved approach to the analysis of Mining Software Repositories (MSR) datasets via metadata enrichment, FAIRness assessment, and topic-driven analysis.
By Aabia Ather, Muhammad Usayd Ather, Qurat-Ul-Ain Somroo, Muhammad Khuram Shahzad
The paper introduces WaterBERT, a domain‑adapted encoder model trained on a 2.97‑billion‑token water treatment corpus to capture domain‑specific semantics for literature mining. Fine‑tuned versions of WaterBERT outperform general‑purpose and other domain BERT models on tasks such as treatment process classification, named entity recognition, and relation extraction. The authors also demonstrate WaterBERT’s utility in large‑scale processing, generating coherent research topics, building a structured knowledge graph from 693,211 abstracts, and creating a Water Knowledge‑Enhanced Retrieval System that surpasses text‑based baselines.
By Mudi Zhai (UNSW Water Research Centre, School of Civil and Environmental Engineering, The University of New South Wales, Sydney, NSW 2052, Australia), Ruihong Qiu (School of Electrical Engineering and Computer Science, The University of Queensland, Brisbane, QLD 4072, Australia), Qingyun Zeng (Microsoft Copilot Studio AI, Redmond, WA 98052, United States, Departments of Mathematics & Department of Computer and Information Science, University of Pennsylvania, Philadelphia, PA 19104, United States), T. David Waite (UNSW Water Research Centre, School of Civil and Environmental Engineering, The University of New South Wales, Sydney, NSW 2052, Australia), Bing-Jie Ni (UNSW Water Research Centre, School of Civil and Environmental Engineering, The University of New South Wales, Sydney, NSW 2052, Australia), Haoran Duan (UNSW Water Research Centre, School of Civil and Environmental Engineering, The University of New South Wales, Sydney, NSW 2052, Australia, Department of Civil Engineering, The University of Hong Kong, Pokfulam, Hong Kong SAR, China)
The paper introduces an ontology‑guided multi‑agent framework for extracting evaluation objects from academic review texts, addressing challenges such as abstractness, context‑dependency, and ambiguous type boundaries. The system combines candidate discovery, ontology‑constrained classification, and domain review, achieving high precision (90.33%) and recall (84.55%) and outperforming rule‑based and zero‑shot baselines. Ablation studies show that the multi‑agent workflow boosts recall and stability, while ontology‑based constraints improve fine‑grained classification and reduce category confusion.
By Haolin Chen, Hongyi Dong, Yu Zhu, Yijia Hong, Leiqing Niu, Jiyuan Ye
arXiv:2609.00228v1 Announce Type: new
Abstract: Scientific domain entity linking (EL) differs from general domain EL because mentions and entity names often lack lexical overlap. Another challenge is...
By Md Rasel Khondokar, Qiao Qiao, Farjana Sultana Samia, Nhat Le, Yuepei Li, Qi Li
arXiv:2607. 27726v1 Announce Type: cross Abstract: Deep research over data lakes requires an LLM agent to investigate evidence across thousands of heterogeneous tables and passages to synthesize a report.
By Dhruv Agarwal, Rishitha Guttapalle Mohan, Aarti Kumari, Ashi Sinha, Athulya Anil, Kavitha Srinivas, Horst Samulowitz, Andrew McCallum
arXiv:2609.16498v1 Announce Type: cross
Abstract: Research data repositories are essential infrastructure for scientific inquiry and for ensuring that datasets follow FAIR (Findable, Accessible, Inte...
By Daniel Ebanks, Devika Jain