arXiv Machine Learning

Private and interpretable clinical prediction with quantum-inspired tensor train models

The paper demonstrates that publicly available clinical machine learning models, such as logistic regression (LR), pose significant privacy risks because attackers can recover model parameters and identify training cohorts through various membership inference attacks. The authors show that even small cohorts can be reliably identified and that common practices like cross-validation can worsen the risk. To mitigate this, they propose a quantum-inspired defense that tensorizes discretized models into tensor trains (TTs), which obfuscates parameters, preserves accuracy, and maintains interpretability while providing black‑box protection comparable to Differential Privacy.

arXiv AI
Jul 23

SynPre-FL: Synthetic data-driven pretraining integrated Federated Learning training framework

arXiv:2607. 19524v1 Announce Type: cross Abstract: Federated learning (FL) offers a promising approach to privacy-preserving clinical risk prediction, but its deployment remains limited by restricted data sharing, client heterogeneity, class imbalance, and the lack of realistic tabular electronic health record (EHR) benchmarks.

By Akarsh K Nair, Muhammad Arifur Rahman, Nicholas Shopland, Andy Burton, Jun He, Yuan Shen, David Baldwin, Emma O'Dowd, Amna Burzic, Mufti Mahmud, David J. Brown
Hugging Face Trending Papers
Jun 3

Federated Learning for Multi-Center Sepsis Early Prediction with Privacy-Preserving

Privacy-sensitive and distributed characteristics of multi-center medical data bring severe obstacles to centralized modeling for accurate early prediction of sepsis. Federated learning (FL) has attracted growing attention as a promising framework for collaborative model development, as it allows multiple institutions to jointly train predictive models without directly sharing or centralizing raw data.

arXiv AI
Jun 18

PSyGenTAB: A Privacy-Preserving Framework for Synthetic Clinical Tabular Data Generation via Constrained Optimization

arXiv:2606. 18518v1 Announce Type: cross Abstract: The development of medical AI is constrained by limited access to high-quality clinical data due to institutional silos and strict privacy regulations such as HIPAA and GDPR.

By Arshia Ilaty, Hossein Shirazi, Manasi Chitale, Kedar Hegde, Dhanalakshmi Ramesh, Rashmi S. Manjunath, Amir Rahmani, Hajar Homayouni
arXiv Machine Learning
Jun 2

Profiling Privacy Preservation Against Gradient Inversion Attacks in Tabular Federated Learning

arXiv:2606. 00986v1 Announce Type: new Abstract: Federated learning (FL) enables multiple data holders to train machine learning models collaboratively without centralizing raw data, making it useful in privacy sensitive domains such as healthcare and institutional data sharing.

By Ivo Osterberg Nilsson, Maximilian Birr Engvall, Viktor Valadi, Teddy Lazebnik
arXiv Machine Learning
Aug 19

Quantifying Memorization and Privacy Risks in Genomic Language Models

The paper introduces a comprehensive privacy evaluation framework for genomic language models (GLMs) that quantifies memorization risks using perplexity-based detection, canary sequence extraction, and membership inference. By planting canary sequences at different repetition rates in synthetic and real datasets, the authors systematically assess how repetition, model capacity, and training dynamics affect memorization across various GLM architectures. The study demonstrates that GLMs do memorize training data to varying degrees and that no single attack method fully captures this risk, highlighting the necessity of multi-vector privacy auditing for genomic AI systems.

By Alexander Nemecek, Wenbiao Li, Xiaoqian Jiang, Jaideep Vaidya, Erman Ayday
arXiv Machine Learning
Aug 27

Are LLM-Enhanced GNNs Privacy-Safe?

The paper evaluates privacy risks in graph neural networks enhanced by large language models (LLMs). Using a five‑stage framework, the authors test six real‑world text‑attributed graph datasets with 42 model configurations and six privacy attack methods across link, label, and membership inference threats. Results show that LLM‑enhanced GNNs are more vulnerable than shallow baselines, with semantic enrichment amplifying exploitable signals, and that differential privacy can reduce risk but at a significant cost to utility.

By Longzhu He, Zelang Wen, Chaozhuo Li, Sen Su
arXiv AI
3d ago

GRIN+: Towards Fast Yet Effective Machine Unlearning for Imbalanced Medical Data

GRIN+ is a new machine unlearning framework that targets fast and precise data erasure in imbalanced medical datasets. It separates unlearning‑specific knowledge from general representations by analyzing gradient contributions of forget and retain sets, introduces a class‑adaptive influence scoring to counter gradient dominance, and uses a direction‑constrained update to protect essential clinical knowledge. Benchmarks on skin cancer, brain tumor, and breast ultrasound data show that GRIN+ balances privacy, efficiency, and utility, achieving high diagnostic accuracy and faster runtime than existing methods.

By Minghui Huang, Junxiao Wang
arXiv Machine Learning
2d ago

Memorisation bias in medical AI

arXiv:2609.17223v1 Announce Type: new Abstract: Medical AI models hold immense potential to improve patient outcomes, but they are also known to unintentionally memorise individual records from their...

By Moritz A. Knolle, Martin J. Menten, Laurin Lux, M\'elanie Roschewitz, Emma A. M. Stanley, Georgios Kaissis, Daniel Rueckert, Ben Glocker