arXiv:2603.08166v2 Announce Type: replace
Abstract: Automated Drug Combination Extraction (DCE) from large-scale biomedical literature is crucial for advancing precision medicine and pharmacological...
By Zhijun Wang, Ling Luo, Dinghao Pan, Huan Zhuang, Lejing Yu, Yuanyuan Sun, Hongfei Lin
arXiv:2609.12544v1 Announce Type: new
Abstract: Clinical de-identification relies on accurately identifying personally identifiable information (PII). However, manually annotated datasets are costly...
By Linh Uyen Le, Christian Hoang, Huy Hoang Ha
arXiv:2606. 13051v1 Announce Type: new Abstract: Despite advances in information extraction driven by deep learning and large language models, performance gaps remain in highly specialized biomedical fields, where domainspecific complexity poses challenges for generalist models.
By Fabien Maury (Imagine - U1163, HeKA | U1346), Sol\`ene Grosdidier (Imagine - U1163), Maud de Dieuleveult (Imagine - U1163), Adrien Coulet (HeKA | U1346)
The paper introduces MiNER, a fine‑tuned biomedical NLP system that uses BioBERT to extract malaria‑related named entities from scientific literature. It builds a large, annotated corpus of malaria articles, preprocesses the text, and applies supervised learning to improve extraction performance. Experiments show that MiNER outperforms other encoding and machine‑learning methods in precision, recall, and accuracy, and the authors release the human‑labeled dataset for further research.
By V. S. Anoop, Devika N
HADRec is a Hierarchy-Aware Drug Recommendation framework that fuses molecular knowledge and electronic health records to improve medication recommendation. It uses LLaMA-7B to encode clinical notes, ChemBERTa to encode drug SMILES strings, and a cross‑attention mechanism for multimodal fusion, while a hierarchical predictor and consistency constraint loss enforce adherence to the ATC classification system. Experiments on MIMIC‑III and MIMIC‑IV show state‑of‑the‑art performance, strong generalization, and well‑calibrated predictions, with counterfactual evaluation indicating clinically aligned reasoning.
By Junke Wang, Hongshun Ling, Li Zhang, Jinjing Wu, Tong Shao, Fang Wang, Yuan Gao
The paper introduces a unified pre‑training framework for medical representations that incorporates hierarchical sub‑token aggregation, partial masking, and cross‑reference mechanisms to better capture the structure of medical codes. The resulting model outperforms existing BERT‑based approaches on pre‑training tasks and downstream clinical predictions, such as dementia onset and hospitalization. An in‑silico drug repositioning study for Alzheimer’s disease demonstrates the framework’s ability to rediscover known drugs and prioritize new hypotheses without external literature, establishing a workflow for hypothesis generation and prioritization based on observational data.
By Yuhei Fujioka, Daitaro Misawa, Shingo Fukuma