Fed-ReMasker is a federated learning approach that adapts the ReMasker masked autoencoder for tabular data imputation, specifically addressing feature-level missingness where entire features are absent at some centers. The method enables centers to impute unobserved features by leveraging knowledge from collaborating institutions. In benchmark tests on synthetic and real-world datasets, Fed-ReMasker achieves the lowest imputation error in the majority of scenarios and remains robust to client heterogeneity, closely matching the performance of a centralized model.
By Ioannis Papathanail, Rooholla Poursoleymani, Lubnaa Abdur Rahman, Stavroula Georgia Mougiakakou
arXiv:2505. 00571v3 Announce Type: replace-cross Abstract: Machine Learning (ML) is gaining popularity in epidemiology and healthcare studies for hypothesis-free discovery of risk and protective factors.
By Giorgio Spadaccini, Marjolein Fokkema, Mark A. van de Wiel
arXiv:2411. 05196v3 Announce Type: replace Abstract: This study presents DhondtXAI as a SHAP-independent, D'Hondt-based attribution framework for tabular XAI.
By Turker Berk Donmez
arXiv:2607. 07725v1 Announce Type: cross Abstract: Genomic prediction models often fail to transfer across institutions because sequencing panels differ across sites, creating structural feature missingness at deployment.
By Muhammet Sami Yavuz, Ayhan Can Erdur, Sabri Mustafa Kahya, Benedikt Wiestler, Jana Lipkova
The paper introduces a missingness‑aware conformal calibration method for mortality prediction that accounts for cross‑hospital distribution shifts. By selecting a measurement on an independent sample, grouping patients by whether that measurement is recorded, and applying Mondrian calibration within each group, the method avoids reusing calibration outcomes. Experiments on eICU and MIMIC‑IV data show that, compared to pooled calibration, it reduces the worst‑group coverage gap by a median of 1.9 percentage points across six settings, though the benefit varies with predictor and hospital.
By Liang You, Dongwen Ou, Hengyu Shi, Siyuan Dai
arXiv:2607. 05613v1 Announce Type: new Abstract: Clinical care often relies on key laboratory indicators, yet real-world patient visits are sparse and tests are ordered irregularly, leading to pervasive missingness.
By Xinrui He, Mengting Ai, Junting Wang, Curtiss B. Cook, Jingrui He
arXiv:2601. 05151v3 Announce Type: replace-cross Abstract: Feature selection (FS) is essential for biomarker discovery and clinical predictive modeling.
By Anastasiia Bakhmach, Paul Dufoss\'e, Simon Charpigny, Florence Monville, Laurent Greillier, Fabrice Barl\'esi, S\'ebastien Benzekry
arXiv:2604. 15107v2 Announce Type: replace-cross Abstract: Shapley values provide a flexible framework for attributing feature contributions to model predictions, but they are not naturally suited for feature selection: a feature may receive a positive attribution even when it is redundant given the remaining variables.
By Chenghui Zheng, Garvesh Raskutti
The study uses large-scale telehealth data and machine learning to classify self‑reported chronic kidney disease (CKD) status and identify key risk factors. A customized stacked ensemble model achieved balanced accuracy of 72.56–76.12% and AUROC of 79.59–82.29%. SHapley Additive exPlanations revealed that regular medical check‑ups, age, blood pressure, and mental health stress indicators are critical predictors of CKD.
By Md. Atik Shams, David Eisenberg, Sumaiya Fatema, Asma Sultana, D. M Hasibul Islam, Junnatul Mawa, Anindita Datta, Nafiya Ahmed, Danastan Tasaouf Mridula, SK. Sazid Mahmud, Simon Bin Akter, Tanjila Helaly, Jorge Fresneda Fernandez, Humayera Islam, Tanmoy Sarkar Pias
arXiv:2609.26142v1 Announce Type: cross
Abstract: Augmented inverse-probability weighting (AIPW), targeted maximum likelihood estimation (TMLE), and double/debiased machine learning (DML) are three r...
By M. Ehsan Karim
arXiv:2605. 25050v2 Announce Type: replace-cross Abstract: Integrating multimodal datasets in clinical oncology is frequently hindered by high dimensionality and blockwise missingness, where entire data sources are unavailable for specific patient subsets.
By Mohamed Boussena, Florence Monville, Jacques Fieschi-Meric, Frederic Vely, Pierre Milpied, Julien Mazieres, Maurice Perol, Eric Vivier, Laurent Greillier, Fabrice Barlesi, Sebastien Benzekry
arXiv:2607. 09165v1 Announce Type: cross Abstract: Achieving early and timely diagnosis and treatment for disease is a major challenge.
By Qingchu Jin, Felistas Mazhude, Jamie B. Rabb, Robert S. Kramer, Douglas B. Sawyer, Raimond L. Winslow