arXiv:2606. 07141v1 Announce Type: cross Abstract: Language models trained for clinical disease inference are trained on patient data, which may include sensitive and private information, and data owners may request the removal of their data from a trained model due to privacy or copyright concerns.
By Anurag Sharma, Sai Teja Chunchu, Prasenjit Mitra, Sandipan Sikdar, Koustav Rudra
arXiv:2603. 15263v2 Announce Type: replace-cross Abstract: Self-supervised learning (SSL) has revolutionized representation learning, with Joint-Embedding Architectures (JEAs) emerging as an effective approach for capturing semantic features.
By Konstantinos Almpanakis, Anna Kreshuk
Fed-ReMasker is a federated learning approach that adapts the ReMasker masked autoencoder for tabular data imputation, specifically addressing feature-level missingness where entire features are absent at some centers. The method enables centers to impute unobserved features by leveraging knowledge from collaborating institutions. In benchmark tests on synthetic and real-world datasets, Fed-ReMasker achieves the lowest imputation error in the majority of scenarios and remains robust to client heterogeneity, closely matching the performance of a centralized model.
By Ioannis Papathanail, Rooholla Poursoleymani, Lubnaa Abdur Rahman, Stavroula Georgia Mougiakakou
arXiv:2606. 19643v1 Announce Type: cross Abstract: Motivated by the privacy, sensitivity and sharing limitations of health data, we present a comprehensive pipeline for inference of Bayesian mixture models within a federated learning setting, i.
By Julie Fendler, Francesca L. Crowe, Tom Marshall, Sylvia Richardson, Paul D. W. Kirk
arXiv:2604. 07085v2 Announce Type: replace Abstract: In electronic health records (EHRs), clustering patients and distinguishing disease subtypes are key tasks to elucidate pathophysiology and aid clinical decision-making.
By Manar D. Samad, Yina Hou, Shrabani Ghosh
SAGE (Subpopulation-Aware Generative Enhancement) is a two-stage generative augmentation framework designed to mitigate spurious correlations in machine learning when group labels are unavailable. It uses cluster-derived sub-labels and class labels to fine‑tune a conditional generative model and text encoder, producing synthetic data that fills underrepresented regions and creates a balanced validation set for last‑layer reweighting. Experiments show SAGE improves worst‑group accuracy to 89.5%, 85.7%, and 79.1% on Waterbirds, CelebA, and MetaShift, outperforming existing group‑label‑free baselines by up to 7.7 percentage points.
By Yiming Luo, Rongqiang Zhao, Jie Liu
arXiv:2608.28923v1 Announce Type: cross
Abstract: Data augmentation is a cornerstone of deep learning pipelines, yet existing strategies treat it as a static, model-agnostic preprocessing step, eithe...
By Noah Videcrantz, Mostafa Mehdipour Ghazi
arXiv:2607. 20641v1 Announce Type: new Abstract: Federated learning (FL) enables multiple clinical institutions to collaboratively train a shared disease classifier without centralizing patient data.
By Afsaneh Mahanipour, Hana Khamfroush
arXiv:2506. 22427v2 Announce Type: replace-cross Abstract: We propose CLoVE (Clustering of Loss Vector Embeddings), a novel algorithm for Clustered Federated Learning (CFL).
By Randeep Bhatia, Nikos Papadis, Murali Kodialam, TV Lakshman, Sayak Chakrabarty
GRIN+ is a new machine unlearning framework that targets fast and precise data erasure in imbalanced medical datasets. It separates unlearning‑specific knowledge from general representations by analyzing gradient contributions of forget and retain sets, introduces a class‑adaptive influence scoring to counter gradient dominance, and uses a direction‑constrained update to protect essential clinical knowledge. Benchmarks on skin cancer, brain tumor, and breast ultrasound data show that GRIN+ balances privacy, efficiency, and utility, achieving high diagnostic accuracy and faster runtime than existing methods.
By Minghui Huang, Junxiao Wang
arXiv:2608. 15310v1 Announce Type: cross Abstract: Multimodal data collected by heterogeneous devices are used for collaborative training, where federated learning (FL) serves as a key paradigm for effective distributed modeling with data privacy preservation.
By Zhenyan Liu, Hua Zhang, Haoran Gao, Qi Li, Hongliang Zhu, Huiyu Zhou, Zongliang Shen, Yanxin Xu, Jiahui Wang
The paper introduces Prototype Purification and Regulation (PPR), a multi‑label few‑shot learning framework for medical image classification that addresses two key limitations of existing metric‑based meta‑learning methods. PPR first purifies prototypes by using sample‑level comorbidity scores to highlight disease‑specific features, then regulates inter‑class prototype distances with disease‑level comorbidity statistics to create a comorbidity‑aware embedding space. Experiments on four chest X‑ray datasets, including cross‑domain tests, show that PPR outperforms state‑of‑the‑art methods, improving disease detection and demonstrating robust generalization and clinical applicability.
By Ying-Chih Lin, Po-Chih Kuo, Yong-Sheng Chen