arXiv:2609.19071v1 Announce Type: new
Abstract: Extracting SNP-phenotype associations from biomedical literature is vital but challenging. We benchmarked diverse NLP models, including MLMs, hybrid ar...
By Claudiu Creanga, Teodor Marchitan, Liviu P. Dinu
The paper introduces MiNER, a fine‑tuned biomedical NLP system that uses BioBERT to extract malaria‑related named entities from scientific literature. It builds a large, annotated corpus of malaria articles, preprocesses the text, and applies supervised learning to improve extraction performance. Experiments show that MiNER outperforms other encoding and machine‑learning methods in precision, recall, and accuracy, and the authors release the human‑labeled dataset for further research.
By V. S. Anoop, Devika N
arXiv:2606. 15412v1 Announce Type: cross Abstract: Biomedical relation extraction (BioRE) is a key step in transforming biomedical literature into structured knowledge.
By Jakob Mraz, Toma\v{z} Curk, Bla\v{z} Zupan
PiPMRE is a new pipeline for medical relation extraction that uses language models instead of traditional tagging schemes. The framework includes a relation generator that produces multiple relational triplets from a text and a relation filter that scores and selects the most reliable triplets. Experiments on two public datasets show that PiPMRE outperforms previous state‑of‑the‑art methods, improving recall by 5.6 points and accuracy by 4.4 points, and it also performs well in few‑shot scenarios.
By Jiaxin Duan, Fengyu Lu, Junfei Liu
arXiv:2506. 02212v2 Announce Type: replace-cross Abstract: Natural Language Processing (NLP) has transformed various fields beyond linguistics by applying techniques originally developed for human language to the analysis of biological sequences.
By Ella Rannon, David Burstein
arXiv:2606. 29639v1 Announce Type: cross Abstract: Automatic prompt optimization is still underexplored for episodic few-shot relation extraction with smaller language models.
By Aunabil Chakma, Mihai Surdeanu, Eduardo Blanco
arXiv:2607. 08803v1 Announce Type: cross Abstract: The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology.
By Hyunjin Seo, Hyeon Hwang, Gyubok Lee, Jay Shin, Jimin Park, Taesoo Kim, Sanghoon Lee, Hongjoon Ahn, Sungjun Han, Sangwon Jung
arXiv:2609.26347v1 Announce Type: cross
Abstract: The scarcity of non-English language data in specialized domains significantly limits the development of effective Natural Language Processing (NLP)...
By Julien Knafou, Luc Mottin, Ana\"is Mottaz, Alexandre Flament, Patrick Ruch
arXiv:2603. 13673v2 Announce Type: replace Abstract: Accurate extraction of Alzheimer's Disease and Related Dementias (ADRD) phenotypes from electronic health records (EHR) is critical for early-stage detection and disease staging.
By Mingchen Shao, Yuzhang Xie, Carl Yang, Jiaying Lu
BELXTR is a new biomedical entity linking model that uses a multi‑vector (late interaction) architecture to preserve token‑level matching information, unlike traditional embedding‑based approaches that compress mentions into a single vector. By extending the XTR model with a task‑specific training objective and active query expansion, BELXTR achieves state‑of‑the‑art performance on half of ten evaluated corpora, with an average 5‑percentage‑point gain in recall@1. The model shows especially strong results on cross‑species gene disambiguation, outperforming an LLM‑powered retrieve‑and‑rerank pipeline and approaching a specialized rule‑based system.
By Samuele Garda, Ulf Leser
arXiv:2606. 13051v1 Announce Type: new Abstract: Despite advances in information extraction driven by deep learning and large language models, performance gaps remain in highly specialized biomedical fields, where domainspecific complexity poses challenges for generalist models.
By Fabien Maury (Imagine - U1163, HeKA | U1346), Sol\`ene Grosdidier (Imagine - U1163), Maud de Dieuleveult (Imagine - U1163), Adrien Coulet (HeKA | U1346)
arXiv:2606. 19852v1 Announce Type: cross Abstract: Information extraction from pathology reports is essential for cancer staging, tumor registry population.
By Aman Pathak, Cheng Peng, Mengxian Lyu, Ziyi Chen, Reema Solan, Sankalp Talankar, Yasir Khan, Hiren Mehta, Aokun Chen, Yi Guo, Yonghui Wu