arXiv:2609.00228v1 Announce Type: new
Abstract: Scientific domain entity linking (EL) differs from general domain EL because mentions and entity names often lack lexical overlap. Another challenge is...
By Md Rasel Khondokar, Qiao Qiao, Farjana Sultana Samia, Nhat Le, Yuepei Li, Qi Li
arXiv:2608. 19201v1 Announce Type: cross Abstract: Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific literature are often inconsistent and difficult to systematically identify at scale.
By Hao Xuan, Rithvij Pasupuleti, Ben Liu, Haishuo Sun, Jun Zhang, Zijun Yao, Cuncong Zhong
arXiv:2608. 08636v1 Announce Type: cross Abstract: Scientific named entity recognition (SciNER) plays a crucial role in information extraction and knowledge discovery from scientific texts.
By Tong Bao, Yi Zhao, Heng Zhang, Chengzhi Zhang
BioELX is a retrieve‑rerank framework for cross‑lingual biomedical entity linking that tackles two key problems: the English‑biased UMLS alias training data and the degradation caused by naïvely adding context. It fine‑tunes SapBERT_multi with Wikidata‑derived cross‑lingual alias supervision to create shared concept neighborhoods, and then reranks candidates using pretrained LLMs with mention‑anchored prompting to focus on the target mention. Experiments demonstrate state‑of‑the‑art performance on four benchmarks, improving Recall@1 by 4.8–18.2 percentage points without task‑specific annotations.
By Yi Wang, Corina Dima, Liangyu Zhong, Steffen Staab
The paper introduces a weakly supervised framework for extracting dataset mentions from forced displacement and Fragile, Conflict, and Violence (FCV) documents. It uses a lightweight model trained on general research literature to generate candidate mentions, which are then refined by a large language model that validates or rejects them and corrects boundaries. The refined annotations are augmented with synthetic and contrastive examples to fine‑tune the model, achieving 74.1% precision and 70.5% recall on a benchmark of 1,706 passages, with higher precision (89.5%) on passages that contain dataset references.
By Rafael Macalaba, Aivin V. Solatorio, Patrick Michael Brock, Olivier Dupriez
arXiv:2609.23231v1 Announce Type: new
Abstract: Cross-lingual information retrieval (CLIR) is increasingly important in multi-national industries, where critical technical evidence may exist in a dif...
By Mahdi Astaraki, Mohammad Khodadad, Reza Namazi, Mohammad Arshi Saloot, Amir Reza Behzad Moghadam, Hamidreza Mahyar, Soheila Samiee
This thesis explores how to select and adapt NLP models for global health literature when annotated data and computational resources are scarce. It compares skip‑gram word2vec models trained on increasingly large specialized corpora with BioWordVec for semantic tag discovery, finding that larger coverage does not always yield more useful domain associations. The study also evaluates convolutional spaCy models versus a RoBERTa transformer for named entity recognition, noting a trade‑off between higher F1 scores and longer inference time, and investigates MiniLM few‑shot versus BART‑MNLI zero‑shot classification for multi‑label topic classification, highlighting practical constraints of inference cost.
"whyItMatters":"The work provides empirical guidance on balancing model accuracy and resource demands for building knowledge systems in low‑resource global health settings."
By Genis Skura, Antoine Geissb\"uhler, Jean-Luc Falcone
The paper introduces Multi-Negative Direct Preference Optimisation (MDPO), a pairwise objective that compares the correct entity with all valid rejected candidates for each mention, extending the single-negative approach used in prior work. MDPO retains the Bradley‑Terry formulation of Direct Preference Optimisation while leveraging the full candidate set through masked, length‑normalised sequence scores. Experiments on French, German, English, Swedish, and Finnish historical newspaper datasets (hipe‑2020 and newseye) show that MDPO outperforms both supervised fine‑tuning and single‑negative DPO, especially for NIL mentions, semantic ambiguity, OCR noise, and historically challenging names, and highlight candidate retrieval as a key bottleneck.
By Tien Nam Nguyen, Emanuela Boros, Ahmed Hamdi, Adam Jatowt, Micka\"el Coustaty, Antoine Doucet
arXiv:2608. 07254v1 Announce Type: cross Abstract: The increasing specialization of scientific research challenges existing classification systems, which provide effective representations of broad disciplines and research topics but often fail to capture the fine-grained conceptual structure of contemporary science.
By Daniele Raimondi, Feichi Lu, Oliver Grun, Mariia Eremina, Andrea Perlato
SciNLP is a new benchmark dataset for full‑text entity and relation extraction in the NLP domain, comprising 60 manually annotated papers with 6,429 entities and 1,649 relations. It is the first dataset to provide full‑text annotations of entities and their relationships specifically for NLP literature. Experiments show that models trained on SciNLP outperform baselines on certain tasks, and the dataset enabled the automatic construction of a fine‑grained knowledge graph with an average node degree of 3.3.
By Decheng Duan, Yingyi Zhang, Jitong Peng, Chengzhi Zhang
arXiv:2609.14770v1 Announce Type: cross
Abstract: Generalisations are common in scientific communication, even though they are semantically ambiguous. An automated method is needed to identify and ca...
By Chenxin Diao, Nataliya Stepanova, Emily Allaway
The paper introduces GPTKB 2.0, a method for building disambiguated knowledge bases directly from large language models. It addresses the lack of native entity representation in LLMs by performing on‑the‑fly disambiguation of entities, relations, and classes, achieving a million‑scale KB with over 1 million disambiguated entities and 38.4 million triples. The authors analyze trade‑offs among accuracy, scale, and cost, and release the system at https://gptkb.org/.
By Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski