arXiv AI By Tong Bao, Yi Zhao, Heng Zhang, Chengzhi Zhang

Enhancing Scientific Named Entity Recognition via Large Language Models: A Type-driven Multi-task Learning Approach

Read the original on arXiv AI →

arXiv:2608. 08636v1 Announce Type: cross Abstract: Scientific named entity recognition (SciNER) plays a crucial role in information extraction and knowledge discovery from scientific texts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 16

SciNLP: A Domain-Specific Benchmark for Full-Text Scientific Entity and Relation Extraction in NLP

SciNLP is a new benchmark dataset for full‑text entity and relation extraction in the NLP domain, comprising 60 manually annotated papers with 6,429 entities and 1,649 relations. It is the first dataset to provide full‑text annotations of entities and their relationships specifically for NLP literature. Experiments show that models trained on SciNLP outperform baselines on certain tasks, and the dataset enabled the automatic construction of a fine‑grained knowledge graph with an average node degree of 3.3.

By Decheng Duan, Yingyi Zhang, Jitong Peng, Chengzhi Zhang
arXiv Computation and Language
Sep 10

Multi-Functional Embedding Models for Funder Name Disambiguation in Scientific Publication Records

The paper introduces a multilingual, multi-functional framework for disambiguating funder names in scientific publications, using a training dataset that merges the Research Organization Registry with Web of Science and Crossref Open Funder Registry data. By applying multi-task learning with contrastive and multiple negatives ranking losses, the authors fine‑tune open‑weight embedding models from the Sentence Transformer, Gemma, and Qwen3 families, achieving over 90% accuracy in matching Web of Science funder names to ROR identifiers and surpassing general‑purpose LLMs by more than 0.1. For funders not present in ROR, a similarity network is constructed to identify clusters, and the study discusses challenges related to smaller and non‑English‑speaking funders.

By Kanyao Han, Zhiwen You, Jinseok Kim, Jana Diesner
arXiv Computation and Language
Sep 22

Custom Named Entity Recognition and Topic Classification for Global Health Publications

This thesis explores how to select and adapt NLP models for global health literature when annotated data and computational resources are scarce. It compares skip‑gram word2vec models trained on increasingly large specialized corpora with BioWordVec for semantic tag discovery, finding that larger coverage does not always yield more useful domain associations. The study also evaluates convolutional spaCy models versus a RoBERTa transformer for named entity recognition, noting a trade‑off between higher F1 scores and longer inference time, and investigates MiniLM few‑shot versus BART‑MNLI zero‑shot classification for multi‑label topic classification, highlighting practical constraints of inference cost. "whyItMatters":"The work provides empirical guidance on balancing model accuracy and resource demands for building knowledge systems in low‑resource global health settings."

By Genis Skura, Antoine Geissb\"uhler, Jean-Luc Falcone
arXiv Computation and Language
Sep 23

BELXTR: Biomedical Entity Linking via Contextualized Token Retrieval

BELXTR is a new biomedical entity linking model that uses a multi‑vector (late interaction) architecture to preserve token‑level matching information, unlike traditional embedding‑based approaches that compress mentions into a single vector. By extending the XTR model with a task‑specific training objective and active query expansion, BELXTR achieves state‑of‑the‑art performance on half of ten evaluated corpora, with an average 5‑percentage‑point gain in recall@1. The model shows especially strong results on cross‑species gene disambiguation, outperforming an LLM‑powered retrieve‑and‑rerank pipeline and approaching a specialized rule‑based system.

By Samuele Garda, Ulf Leser