The paper introduces a web‑data curation recipe for pretraining medical encoders, addressing the scarcity of large, diverse corpora in dense‑terminology domains like medicine. It proposes two complementary techniques: medical‑term density filtering to select documents rich in medical terminology, and signal‑amplifying rephrasing that uses an LLM to rewrite documents into denser variants with broader entity contexts. Applied to French medical NLP, the recipe produces the FineMed corpus and the DoctoBERT encoder family, achieving state‑of‑the‑art results on the DrBenchmark public benchmark and a proprietary clinical NER task.
By Bofeng Huang, Jacques Sun, Diane Bouchacourt, Nicolas Barascud, Fajwel Fogel
The paper introduces Hyperbolic Clinical Ontology Embeddings (HCOE), a method that transforms frozen BioBERT embeddings into a Poincaré ball to capture medical code hierarchies. HCOE employs ontology-guided contrastive learning and coarse‑to‑fine ontology‑path aggregation, leveraging ICD, CCS, and ATC hierarchies. Experiments on MIMIC‑IV demonstrate superior performance in clinical relation prediction, hierarchy transfer, and various predictive tasks such as mortality, readmission, medication recommendation, and rare drug prediction.
By Yixuan Li, Weihao Li, Ziyang Song
Biomedical ontology normalization maps free-text expressions to standardized concepts, enabling consistent integration and analysis of biomedical data. This task remains challenging because lexical va...
The paper proposes a behavior-based fusion model that combines large language models (LLMs) with ontology rankers to improve rare-disease diagnosis. By examining ranked lists, agreement, and ontology support, the model learns how much to rely on each system per case, achieving significant recall gains on Phenopacket Store and RAMEDIS benchmarks. Importantly, the fused diagnoses retain candidate-level ontology evidence for inspection.
By Zhaoyang Jiang, Zhizhong Fu, Yunsoo Kim, Zicheng Li, Xuanqi Peng, Fei Teng, Jiacong Mi, Honghan Wu
arXiv:2609.10055v1 Announce Type: cross
Abstract: Biomedical ontology normalization maps free-text expressions to standardized concepts, enabling consistent integration and analysis of biomedical dat...
By Jie Song, Zhichuan Xu, Ziyu Lu, Meng Xiao, Cheng Bi, Yuxin Zhang, Xin Zheng, Xiaoran Li, Qiongfang Cao, Hao Yang, Bairong Shen
arXiv:2607. 01977v1 Announce Type: new Abstract: Ontology learning (OL) aims to automatically construct structured knowledge models from text, yet progress remains fragmented across methods, domains, and evaluation practices.
By Hamed Babaei Giglou, Jennifer D'Souza, Andrei Aioanei, Nandana Mihindukulasooriya, S\"oren Auer
arXiv:2604. 16878v2 Announce Type: replace Abstract: Early prediction of severe clinical deterioration and remaining length of stay can enable timely intervention and better resource allocation in high-acuity settings such as the ICU.
By Zhongyuan Liang, Junhyung Jo, Hyang-Jung Lee, Sang Kyu Kim, Irene Y. Chen
arXiv:2606. 19266v1 Announce Type: cross Abstract: The development of large language models (LLMs) has led to an increased focus on their adaptation to specialized domains and languages, yet the effectiveness of domain adaptation strategies remains unclear.
By Ikram Belmadani, Oumaima El Khettari, Carlos Ramisch, Frederic Bechet, Richard Dufour, Benoit Favre
arXiv:2602.17826v2 Announce Type: replace
Abstract: Language models exhibit fundamental limitations -- hallucination, brittleness, and lack of formal grounding -- that are particularly problematic in...
By Marcelo Labre
arXiv:2608.31118v1 Announce Type: new
Abstract: The effect of Large Language Model (LLM) scale on ontology learning (OL) performance remains insufficiently characterized. We present a controlled eval...
By Hamed Babaei Giglou, S\"oren Auer, Jennifer D'Souza
arXiv:2607. 08803v1 Announce Type: cross Abstract: The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology.
By Hyunjin Seo, Hyeon Hwang, Gyubok Lee, Jay Shin, Jimin Park, Taesoo Kim, Sanghoon Lee, Hongjoon Ahn, Sungjun Han, Sangwon Jung
The paper introduces UdonCare, a hierarchy‑pruning method that iteratively partitions patients into latent domains using medical ontologies, aiming to improve domain generalization in clinical prediction tasks. It addresses challenges of missing domain labels and lack of clinical insight by discovering hierarchy‑grounded patient domains. Experiments on MIMIC‑III, MIMIC‑IV, and eICU datasets show UdonCare outperforms eight baseline methods across four prediction tasks with significant domain gaps.
By Pengfei Hu, Xiaoxue Han, Fei Wang, Yue Ning