arXiv:2603. 13673v2 Announce Type: replace Abstract: Accurate extraction of Alzheimer's Disease and Related Dementias (ADRD) phenotypes from electronic health records (EHR) is critical for early-stage detection and disease staging.
By Mingchen Shao, Yuzhang Xie, Carl Yang, Jiaying Lu
arXiv:2609.19071v1 Announce Type: new
Abstract: Extracting SNP-phenotype associations from biomedical literature is vital but challenging. We benchmarked diverse NLP models, including MLMs, hybrid ar...
By Claudiu Creanga, Teodor Marchitan, Liviu P. Dinu
The paper introduces pre‑trained models for extracting variant‑phenotype relations from biomedical text, focusing on the SNPPhenA corpus. Fine‑tuning small BERT‑based models, especially DeBERTa, achieves performance close to the current state‑of‑the‑art. Moreover, careful fine‑tuning of Google’s Gemini Pro 1.0 surpasses existing benchmarks on both sentence‑level and abstract‑level relation extraction tasks.
By Claudiu Creanga, Liviu P. Dinu, Daniela Gifu
arXiv:2609.12544v1 Announce Type: new
Abstract: Clinical de-identification relies on accurately identifying personally identifiable information (PII). However, manually annotated datasets are costly...
By Linh Uyen Le, Christian Hoang, Huy Hoang Ha
arXiv:2606. 00031v1 Announce Type: cross Abstract: Coronary artery disease (CAD) remains one of the leading causes of death globally, highlighting the need for reliable predictive systems to support early diagnosis and risk assessment.
By Jeba Maliha, Md Rafiul Kabir
The paper introduces MiNER, a fine‑tuned biomedical NLP system that uses BioBERT to extract malaria‑related named entities from scientific literature. It builds a large, annotated corpus of malaria articles, preprocesses the text, and applies supervised learning to improve extraction performance. Experiments show that MiNER outperforms other encoding and machine‑learning methods in precision, recall, and accuracy, and the authors release the human‑labeled dataset for further research.
By V. S. Anoop, Devika N
arXiv:2606. 15412v1 Announce Type: cross Abstract: Biomedical relation extraction (BioRE) is a key step in transforming biomedical literature into structured knowledge.
By Jakob Mraz, Toma\v{z} Curk, Bla\v{z} Zupan
This thesis explores how to select and adapt NLP models for global health literature when annotated data and computational resources are scarce. It compares skip‑gram word2vec models trained on increasingly large specialized corpora with BioWordVec for semantic tag discovery, finding that larger coverage does not always yield more useful domain associations. The study also evaluates convolutional spaCy models versus a RoBERTa transformer for named entity recognition, noting a trade‑off between higher F1 scores and longer inference time, and investigates MiniLM few‑shot versus BART‑MNLI zero‑shot classification for multi‑label topic classification, highlighting practical constraints of inference cost.
"whyItMatters":"The work provides empirical guidance on balancing model accuracy and resource demands for building knowledge systems in low‑resource global health settings."
By Genis Skura, Antoine Geissb\"uhler, Jean-Luc Falcone
The paper introduces a configurable semantic chunking framework for biomedical information extraction in retrieval‑augmented generation systems. It replaces the fixed‑size chunking stage of BioMedRAG with entity‑preserving windows, trigger‑centered chunking, proposition‑first extraction, tiered trigger prioritization, and hierarchical relation resolution, while keeping the rest of the pipeline unchanged. Experiments on relation extraction benchmarks (GM‑CIHT, DDI, ChemProt) and adverse event classification (ADE) show that the hybrid configuration boosts performance on datasets with explicit relation cues, achieving 82.6% F1 on GM‑CIHT compared to 74.2% with the baseline.
By Riya Ahuja (Institute of Data Science in Biomedicine, TU Braunschweig, Braunschweig, Germany, Braunschweig Integrated Centre of Systems Biology, TU Braunschweig, Braunschweig, Germany), Tim Kacprowski (Institute of Data Science in Biomedicine, TU Braunschweig, Braunschweig, Germany, Braunschweig Integrated Centre of Systems Biology, TU Braunschweig, Braunschweig, Germany), Roya Shiasi Sardoabi (Institute of Data Science in Biomedicine, TU Braunschweig, Braunschweig, Germany, Braunschweig Integrated Centre of Systems Biology, TU Braunschweig, Braunschweig, Germany)
arXiv:2606. 08311v1 Announce Type: new Abstract: Electronic health record (EHR) notes are dense medical documents containing large amounts of information, often filled with complex medical jargon.
By Mahshad Koohi Habibi Dehkordi, Shuxin Zhou, Yehoshua Perl, Fadi P. Deek, James Geller, Gai Elhanan, Andrew J. Einstein, Luke Lindemann, Vipina K. Keloth
arXiv:2601.16753v2 Announce Type: replace-cross
Abstract: Longitudinal information in radiology reports refers to the sequential tracking of findings across multiple examinations over time, which is...
By Xinyi Wang, Grazziela Figueredo, Ruizhe Li, Xin Chen
The article presents a new semantic model for representing scientific evidence, specifically tailored to genetics, that extends existing standards by adding fine‑grained, domain‑specific structure. It aligns with FHIR Evidence and SEPIO, incorporates a compact vocabulary validated by SHACL, and was tested in a human‑AI annotation pilot on six genetics papers, producing 28 evidence items and 95 source‑anchored assertions. The authors argue that this model advances trustworthy, AI‑ready infrastructure for variant interpretation by providing a reference data model and validation schema for genetic evidence.
By Michael Bouzinier, Dmitry Etin