arXiv Machine Learning

MetaboLLM: a metabolomics-specialized large language model for biochemical knowledge integration and predictive metabolite graph construction

arXiv:2608. 06253v1 Announce Type: new Abstract: Metabolomics knowledge is distributed across heterogeneous resources and remains difficult to translate into predictive representations.

arXiv AI
Sep 3

Unifying biomedical knowledge in a modern multimodal graph

OptimusKG is a multimodal biomedical labeled property graph that integrates structured and semi‑structured resources to preserve detailed, type‑specific metadata across molecular, anatomical, clinical, and environmental domains. The graph contains nearly 191,000 nodes, over 21.8 million edges, and more than 67 million property instances derived from 18 ontologies, with a top‑level schema that enforces node and edge constraints while retaining granular provenance. Validation using the PaperQA3 agent found that 70.0% of sampled edges are supported by literature evidence, and the graph offers a standardized resource for machine learning, knowledge‑grounded retrieval, and hypothesis generation in biomedical research.

By Lucas Vittor, Ayush Noori, I\~naki Arango, Joaqu\'in Polonuer, Sam Rodriques, Andrew White, David A. Clifton, Marinka Zitnik
arXiv AI
Sep 7

Hakken: Predicting future discoveries to fill the gaps in today's knowledge

Hakken is a domain‑agnostic system that predicts and explains future scientific discoveries by combining transformer‑based models trained on temporal knowledge graphs with large language model semantic knowledge. It identifies novel relationships between scientific concepts that extend beyond the deductive hull of existing knowledge and provides explanations to help scientists assess these predictions. In the biomedical domain, Hakken set a new benchmark for time‑aware multi‑label relation prediction, generated 1.5 million high‑confidence hypotheses about aging, and experimentally confirmed two predictions that revealed previously undocumented interactions relevant to drug discovery.

By Tarek R. Besold, Uchenna Akujuobi, Pablo Sanchez, Alessandra Toniato, Kana Maruyama, Jihun Choi, Samy Badreddine, Frederick Gifford, Daniel Evans-Yamamoto, Sucheendra K. Palaniappan, Miquel Ferrer, Kae Nagano, Iris Rossell, Tom Joy, Hatem ElShazly, Chrysa Iliopoulou, Christoph Wehner, Thiviyan Thanapalasingam, Susana Nunes, Pedro G. Cotovio, Peter Wurman, Peter Stone, Hiroaki Kitano, Michael Spranger
arXiv AI
Aug 24

LingShu: A Large-Scale Symptom-Centric Contextualized Knowledge Graph Bridging Traditional Chinese Medicine and Modern Biomedicine

LingShu is a large-scale, symptom‑centric knowledge graph that bridges Traditional Chinese Medicine (TCM) and modern biomedicine. It contains 17.33 million entity records and 39.47 million relation records, combining 17.19 million semantic triples with 22.29 million contextualized quadruples to encode conditional medical associations. The graph integrates data from electronic medical records, TCM texts, biomedical ontologies, and curated knowledge bases, and is supported by a web platform offering visualization, reasoning, and evidence‑grounded question answering.

By Rui Hua, Zixin Shu, Kai Chang, Dengying Yan, Jianan Xia, Hui Zhu, Shujie Song, Shurui Yang, Tongxin Wang, Yue Yin, Yu Wei, Lijuan Pei, Yunhui Hu, Hao Xu, Mingzhong Xiao, Xiaodong Li, Haibin Yu, Runshun Zhang, Wenjia Wang, Baoyan Liu, Xuezhong Zhou
arXiv AI
Sep 7

REFINE: LLM Refinement over Budgeted Text-Attributed Graphs for Personalized Medical Concept Representation

REFINE is a framework that refines medical concept representations by creating patient‑specific temporal graphs from a global text‑attributed knowledge graph. It uses a reinforcement learning policy to allocate a personalized graph expansion budget for each observed code, then processes the resulting graph with a heterogeneous GNN and a frozen LLM that refines representations via graph‑aware soft prompts. Experiments on MIMIC‑III and MIMIC‑IV demonstrate that REFINE consistently improves various EHR prediction backbones, surpasses strong baselines, and shows robust gains across ablation studies, KG selection, and data insufficiency scenarios.

By Mohsen Nayebi Kerdabadi, Arya Hadizadeh Moghaddam, Dongjie Wang, Zijun Yao
arXiv Computation and Language
Aug 25

Aligning Biomedical Texts and Knowledge Graphs: A Systematic Comparison of Lightweight Alignment Strategies

The paper introduces a unified framework for aligning biomedical text with knowledge graphs using a lightweight projection learned via contrastive learning, keeping the text encoder and KG embedding model frozen. It evaluates six design choices—text encoder, KG embedding, projection head, triple composition, training direction, and hard‑negative sampling—on a newly created CTD‑Align corpus of 22K chemical‑gene interaction pairs linked to PubMed passages. The study finds that triple composition and training direction have the largest impact, while simpler linear projections over concatenated subject, predicate, and object embeddings yield the best performance.

By Artem Bisliouk, Elizaveta Nosova, Heiko Paulheim, Andreea Iana, Rita T. Sousa
arXiv AI
Jul 7

Predicting Therapeutic Outcome via Aligning Patient-Specific Knowledge Graph and Gene-Level Perturbation Representations

arXiv:2607. 04557v1 Announce Type: cross Abstract: Accurate prediction of patient-specific therapeutic response from pre-treatment transcriptomes is hindered by the scarcity of matched clinical response labels and post-treatment molecular profiles.

By Dongmin Bang, Sugyun An, Inyoung Sung, Ilho Yun, Sun Kim, Sangseon Lee