arXiv AI

OntoBook: Ontology-Grounded Synthetic Textbooks for Medical Encoder Pretraining

arXiv:2607. 18927v1 Announce Type: new Abstract: We present OntoBook, a method that converts medical ontology structure into pretraining signal for encoder language models.

arXiv Machine Learning
Sep 10

Where Does the Signal Live? A Web Data Recipe for Medical Encoder Pretraining

The paper introduces a web‑data curation recipe for pretraining medical encoders, addressing the scarcity of large, diverse corpora in dense‑terminology domains like medicine. It proposes two complementary techniques: medical‑term density filtering to select documents rich in medical terminology, and signal‑amplifying rephrasing that uses an LLM to rewrite documents into denser variants with broader entity contexts. Applied to French medical NLP, the recipe produces the FineMed corpus and the DoctoBERT encoder family, achieving state‑of‑the‑art results on the DrBenchmark public benchmark and a proprietary clinical NER task.

By Bofeng Huang, Jacques Sun, Diane Bouchacourt, Nicolas Barascud, Fajwel Fogel
arXiv AI
6d ago

HCOE: Hyperbolic Clinical Ontology Embeddings from Biomedical Language Models

The paper introduces Hyperbolic Clinical Ontology Embeddings (HCOE), a method that transforms frozen BioBERT embeddings into a Poincaré ball to capture medical code hierarchies. HCOE employs ontology-guided contrastive learning and coarse‑to‑fine ontology‑path aggregation, leveraging ICD, CCS, and ATC hierarchies. Experiments on MIMIC‑IV demonstrate superior performance in clinical relation prediction, hierarchy transfer, and various predictive tasks such as mortality, readmission, medication recommendation, and rare drug prediction.

By Yixuan Li, Weihao Li, Ziyang Song
arXiv Computation and Language
Sep 3

Learning to Fuse LLMs with Ontology Rankers for Rare-Disease Diagnosis

The paper proposes a behavior-based fusion model that combines large language models (LLMs) with ontology rankers to improve rare-disease diagnosis. By examining ranked lists, agreement, and ontology support, the model learns how much to rely on each system per case, achieving significant recall gains on Phenopacket Store and RAMEDIS benchmarks. Importantly, the fused diagnoses retain candidate-level ontology evidence for inspection.

By Zhaoyang Jiang, Zhizhong Fu, Yunsoo Kim, Zicheng Li, Xuanqi Peng, Fei Teng, Jiacong Mi, Honghan Wu
arXiv Computation and Language
Sep 10

OntologyAligner: Ontology-Aligned Retrieval and Hierarchy-Guided Large Language Model Reranking for Biomedical Ontology Normalization

arXiv:2609.10055v1 Announce Type: cross Abstract: Biomedical ontology normalization maps free-text expressions to standardized concepts, enabling consistent integration and analysis of biomedical dat...

By Jie Song, Zhichuan Xu, Ziyu Lu, Meng Xiao, Cheng Bi, Yuxin Zhang, Xin Zheng, Xiaoran Li, Qiongfang Cao, Hao Yang, Bairong Shen
arXiv Machine Learning
Jul 7

OC-Distill: Ontology-aware Contrastive Learning with Cross-Modal Distillation for ICU Risk Prediction

arXiv:2604. 16878v2 Announce Type: replace Abstract: Early prediction of severe clinical deterioration and remaining length of stay can enable timely intervention and better resource allocation in high-acuity settings such as the ICU.

By Zhongyuan Liang, Junhyung Jo, Hyang-Jung Lee, Sang Kyu Kim, Irene Y. Chen
arXiv AI
Jul 13

TheBioCollection: Unified Pre-Training Scale LLM Corpus for Biology

arXiv:2607. 08803v1 Announce Type: cross Abstract: The push toward large language models for biology (BioLM) has created a need for training corpora that can endow models with a genuine understanding of biology.

By Hyunjin Seo, Hyeon Hwang, Gyubok Lee, Jay Shin, Jimin Park, Taesoo Kim, Sanghoon Lee, Hongjoon Ahn, Sungjun Han, Sangwon Jung
arXiv AI
Sep 15

Discovering Hierarchy-Grounded Domains with Adaptive Granularity for Clinical Domain Generalization

The paper introduces UdonCare, a hierarchy‑pruning method that iteratively partitions patients into latent domains using medical ontologies, aiming to improve domain generalization in clinical prediction tasks. It addresses challenges of missing domain labels and lack of clinical insight by discovering hierarchy‑grounded patient domains. Experiments on MIMIC‑III, MIMIC‑IV, and eICU datasets show UdonCare outperforms eight baseline methods across four prediction tasks with significant domain gaps.

By Pengfei Hu, Xiaoxue Han, Fei Wang, Yue Ning