arXiv AI

Cross-Domain Feature Expansion for Tabular Medical Data via Knowledge Graphs Injection

arXiv:2606. 31171v1 Announce Type: new Abstract: Acquiring comprehensive cross-domain biomedical profiles is often costly and time-consuming, resulting in severe data scarcity in medical research.

arXiv AI
Aug 28

Methodological and Conceptual Framework for 5D Multi-Table Analysis: A Unified Approach for Complex Data Reuse

The paper introduces the Relational Hypergraph Transformer (RHT), a unified architecture that models relational databases as hypergraphs and learns pentadimensional embeddings (PentE). RHT applies sparse relational attention whose complexity scales with the average relational degree, making it computationally efficient for large, high‑dimensional, and high‑cardinality datasets. Experiments on the Synthea synthetic electronic health record dataset show that RHT produces more semantically coherent embeddings than tabular, relational, and temporal graph baselines, while remaining scalable, and the authors provide an open‑source implementation and plan clinical validation on MIMIC‑IV.

By Edouard Lansiaux, Hugo Kazzi, Aur\'elien Loison, Slim Hammadi, Emmanuel Chazard
arXiv AI
Sep 3

General Demographic Pre-trained Models for Enhancing Predictive Performance Across Diseases and Population

The paper introduces the General Demographic Pre-trained (GDP) model, a lightweight foundation model that learns representations from the two most common clinical attributes—age and sex. By optimizing encoding and visit‑reordering strategies, GDP embeddings are shown to improve predictive performance when concatenated with raw features across various disease and geographic cohorts. The model outperforms several state‑of‑the‑art tabular foundation models and tree‑based algorithms, demonstrating that enriched demographic embeddings can enhance classification tasks while remaining fully compatible with standard classifiers.

By Li-Chin Chen, Ji-Tian Sheu, Yuh-Jue Chuang
arXiv Machine Learning
Sep 22

From Latent Biomarkers to Clinical Rules: Embedding-Guided Rule Mining and Attribution-Based Translation for Interpretable Tabular Learning

The paper introduces a four-step pipeline that mines decision rules in the latent space of an FT-Transformer and then translates those rules back into measurable clinical features. By treating embedding dimensions that separate patient groups as latent biomarkers, small decision trees are used to extract rules, which are then mapped to raw features using gradient-input saliency and CLS attention attribution. Across six public clinical datasets, the translated rules generally outperformed raw-feature rules, achieving significant AUROC gains, though some high-performing latent rules could not be fully captured by simple raw-feature conditions.

By Majid Lotfian Delouee, Hamed Ayoobi, Sjors G. J. G. In 't Veld, Martijn C. Schut
arXiv AI
Sep 24

Fed-ReMasker: Federated Tabular Imputation under Feature-Level Missingness

Fed-ReMasker is a federated learning approach that adapts the ReMasker masked autoencoder for tabular data imputation, specifically addressing feature-level missingness where entire features are absent at some centers. The method enables centers to impute unobserved features by leveraging knowledge from collaborating institutions. In benchmark tests on synthetic and real-world datasets, Fed-ReMasker achieves the lowest imputation error in the majority of scenarios and remains robust to client heterogeneity, closely matching the performance of a centralized model.

By Ioannis Papathanail, Rooholla Poursoleymani, Lubnaa Abdur Rahman, Stavroula Georgia Mougiakakou
arXiv AI
Sep 7

REFINE: LLM Refinement over Budgeted Text-Attributed Graphs for Personalized Medical Concept Representation

REFINE is a framework that refines medical concept representations by creating patient‑specific temporal graphs from a global text‑attributed knowledge graph. It uses a reinforcement learning policy to allocate a personalized graph expansion budget for each observed code, then processes the resulting graph with a heterogeneous GNN and a frozen LLM that refines representations via graph‑aware soft prompts. Experiments on MIMIC‑III and MIMIC‑IV demonstrate that REFINE consistently improves various EHR prediction backbones, surpasses strong baselines, and shows robust gains across ablation studies, KG selection, and data insufficiency scenarios.

By Mohsen Nayebi Kerdabadi, Arya Hadizadeh Moghaddam, Dongjie Wang, Zijun Yao
arXiv AI
Sep 15

Discovering Hierarchy-Grounded Domains with Adaptive Granularity for Clinical Domain Generalization

The paper introduces UdonCare, a hierarchy‑pruning method that iteratively partitions patients into latent domains using medical ontologies, aiming to improve domain generalization in clinical prediction tasks. It addresses challenges of missing domain labels and lack of clinical insight by discovering hierarchy‑grounded patient domains. Experiments on MIMIC‑III, MIMIC‑IV, and eICU datasets show UdonCare outperforms eight baseline methods across four prediction tasks with significant domain gaps.

By Pengfei Hu, Xiaoxue Han, Fei Wang, Yue Ning
arXiv AI
Jun 2

Retrieval-aligned Tabular Foundation Models Enable Robust Clinical Risk Prediction in Electronic Health Records Under Real-world Constraints

arXiv:2604. 01841v2 Announce Type: replace Abstract: Clinical prediction from structured electronic health records (EHRs) is challenging due to high dimensionality, heterogeneity, class imbalance, and distribution shift.

By Minh-Khoi Pham, Thang-Long Nguyen Ho, Thao Thi Phuong Dao, Tai Tan Mai, Minh-Triet Tran, Marie E. Ward, Una Geary, Rob Brennan, Nick McDonald, Martin Crane, Marija Bezbradica