arXiv AI By Rian Touchent (ALMAnaCH), \'Eric de la Clergerie (ALMAnaCH)

OntoBook: Ontology-Grounded Synthetic Textbooks for Medical Encoder Pretraining

Read the original on arXiv AI →

arXiv:2607. 18927v1 Announce Type: new Abstract: We present OntoBook, a method that converts medical ontology structure into pretraining signal for encoder language models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 10

Where Does the Signal Live? A Web Data Recipe for Medical Encoder Pretraining

The paper introduces a web‑data curation recipe for pretraining medical encoders, addressing the scarcity of large, diverse corpora in dense‑terminology domains like medicine. It proposes two complementary techniques: medical‑term density filtering to select documents rich in medical terminology, and signal‑amplifying rephrasing that uses an LLM to rewrite documents into denser variants with broader entity contexts. Applied to French medical NLP, the recipe produces the FineMed corpus and the DoctoBERT encoder family, achieving state‑of‑the‑art results on the DrBenchmark public benchmark and a proprietary clinical NER task.

By Bofeng Huang, Jacques Sun, Diane Bouchacourt, Nicolas Barascud, Fajwel Fogel
arXiv AI
6d ago

HCOE: Hyperbolic Clinical Ontology Embeddings from Biomedical Language Models

The paper introduces Hyperbolic Clinical Ontology Embeddings (HCOE), a method that transforms frozen BioBERT embeddings into a Poincaré ball to capture medical code hierarchies. HCOE employs ontology-guided contrastive learning and coarse‑to‑fine ontology‑path aggregation, leveraging ICD, CCS, and ATC hierarchies. Experiments on MIMIC‑IV demonstrate superior performance in clinical relation prediction, hierarchy transfer, and various predictive tasks such as mortality, readmission, medication recommendation, and rare drug prediction.

By Yixuan Li, Weihao Li, Ziyang Song
arXiv Computation and Language
Sep 3

Learning to Fuse LLMs with Ontology Rankers for Rare-Disease Diagnosis

The paper proposes a behavior-based fusion model that combines large language models (LLMs) with ontology rankers to improve rare-disease diagnosis. By examining ranked lists, agreement, and ontology support, the model learns how much to rely on each system per case, achieving significant recall gains on Phenopacket Store and RAMEDIS benchmarks. Importantly, the fused diagnoses retain candidate-level ontology evidence for inspection.

By Zhaoyang Jiang, Zhizhong Fu, Yunsoo Kim, Zicheng Li, Xuanqi Peng, Fei Teng, Jiacong Mi, Honghan Wu
arXiv Computation and Language
Sep 10

OntologyAligner: Ontology-Aligned Retrieval and Hierarchy-Guided Large Language Model Reranking for Biomedical Ontology Normalization

arXiv:2609.10055v1 Announce Type: cross Abstract: Biomedical ontology normalization maps free-text expressions to standardized concepts, enabling consistent integration and analysis of biomedical dat...

By Jie Song, Zhichuan Xu, Ziyu Lu, Meng Xiao, Cheng Bi, Yuxin Zhang, Xin Zheng, Xiaoran Li, Qiongfang Cao, Hao Yang, Bairong Shen