arXiv Machine Learning

Binary Gaussian Copula Synthesis: an LLM-powered data augmentation framework for early dialysis prediction in chronic kidney disease

arXiv:2403. 00965v2 Announce Type: replace-cross Abstract: Only a small fraction of patients with chronic kidney disease (CKD) progress to dialysis, creating severe class imbalance that limits the performance of machine learning models for early dialysis prediction.

arXiv Machine Learning
Aug 19

Population Health-Based Machine Learning Reveals Associations Between Psychosocial Factors and Chronic Kidney Disease

The study uses large-scale telehealth data and machine learning to classify self‑reported chronic kidney disease (CKD) status and identify key risk factors. A customized stacked ensemble model achieved balanced accuracy of 72.56–76.12% and AUROC of 79.59–82.29%. SHapley Additive exPlanations revealed that regular medical check‑ups, age, blood pressure, and mental health stress indicators are critical predictors of CKD.

By Md. Atik Shams, David Eisenberg, Sumaiya Fatema, Asma Sultana, D. M Hasibul Islam, Junnatul Mawa, Anindita Datta, Nafiya Ahmed, Danastan Tasaouf Mridula, SK. Sazid Mahmud, Simon Bin Akter, Tanjila Helaly, Jorge Fresneda Fernandez, Humayera Islam, Tanmoy Sarkar Pias
arXiv Machine Learning
Aug 14

CoMedBench: A Multi-Source Benchmark of Synthetic Medical Data Fidelity and Downstream Utility

arXiv:2608. 12805v1 Announce Type: new Abstract: Access to clinical data is essential for developing reliable healthcare machine learning systems, but direct use of electronic health records is constrained by privacy regulation, institutional review, data-use agreements, and the risk of re-identification.

By Akanta Das, Al Amin Farhad, Mrinmoy Sarkar Anto, David Rehkopf, Ayin Vala, Tanmoy Sarkar Pias
arXiv Machine Learning
Aug 4

Development and Validation of a Dynamic Kidney Failure Prediction Model based on Deep Learning: A Real-World Study with External Validation

arXiv:2501. 16388v3 Announce Type: replace Abstract: Background: Chronic kidney disease (CKD), a progressive disease with high morbidity and mortality, has become a significant global public health problem.

By Jingying Ma, Jinwei Wang, Lanlan Lu, Zhiqin Jiang, Mengling Feng, Feifei Zhang, Peng Shen, Yexiang Sun, Shenda Hong, Luxia Zhang
arXiv AI
Jul 23

SynPre-FL: Synthetic data-driven pretraining integrated Federated Learning training framework

arXiv:2607. 19524v1 Announce Type: cross Abstract: Federated learning (FL) offers a promising approach to privacy-preserving clinical risk prediction, but its deployment remains limited by restricted data sharing, client heterogeneity, class imbalance, and the lack of realistic tabular electronic health record (EHR) benchmarks.

By Akarsh K Nair, Muhammad Arifur Rahman, Nicholas Shopland, Andy Burton, Jun He, Yuan Shen, David Baldwin, Emma O'Dowd, Amna Burzic, Mufti Mahmud, David J. Brown
arXiv Machine Learning
Jun 2

Early Prediction of Liver Cirrhosis Up to Two Years in Advance: A Machine Learning Study Benchmarking Against the FIB-4 and APRI Scores

arXiv:2601. 00175v2 Announce Type: replace Abstract: Objective: Develop and evaluate machine learning (ML) models for predicting incident liver cirrhosis (LC) one and two years prior to diagnosis using routinely collected electronic health record (EHR) data and benchmark their performance against the FIB-4 and APRI clinical scores.

By Zhuqi Miao, Ahmed G Qasem, Sujan Ravi, Jason T. Cheng, Abdulaziz Ahmed, Courtney W. Houchen, Sumayah Abed, Dilorom Azimdjanovna Zuparova, Abdulaziz Ahmed
arXiv Machine Learning
Sep 16

Copula Adapted Directed Acyclic Graph for Cluster Representation of Biomedical Data

The paper presents Copula Adapted Directed Acyclic Graph (CopDAG), a framework that combines copula models with an ensemble of causal structure discovery methods based on Directed Acyclic Graphs to represent biomedical data. By capturing non‑Gaussian, non‑linear dependencies and stable causal relationships, CopDAG enables clustering of unlabeled biomedical data using K‑means. Across 16 biomedical datasets, CopDAG achieves the highest normalized clustering accuracy and adjusted Rand index among 12 evaluated methods, and it can predict class labels and provide explainable causal visualizations without relying on data annotations.

By Heranga K. Rathnasekara, Norou Diawara, Manar D. Samad
arXiv AI
Sep 3

General Demographic Pre-trained Models for Enhancing Predictive Performance Across Diseases and Population

The paper introduces the General Demographic Pre-trained (GDP) model, a lightweight foundation model that learns representations from the two most common clinical attributes—age and sex. By optimizing encoding and visit‑reordering strategies, GDP embeddings are shown to improve predictive performance when concatenated with raw features across various disease and geographic cohorts. The model outperforms several state‑of‑the‑art tabular foundation models and tree‑based algorithms, demonstrating that enriched demographic embeddings can enhance classification tasks while remaining fully compatible with standard classifiers.

By Li-Chin Chen, Ji-Tian Sheu, Yuh-Jue Chuang