arXiv AI

PSyGenTAB: A Privacy-Preserving Framework for Synthetic Clinical Tabular Data Generation via Constrained Optimization

arXiv:2606. 18518v1 Announce Type: cross Abstract: The development of medical AI is constrained by limited access to high-quality clinical data due to institutional silos and strict privacy regulations such as HIPAA and GDPR.

arXiv AI
Jul 23

SynPre-FL: Synthetic data-driven pretraining integrated Federated Learning training framework

arXiv:2607. 19524v1 Announce Type: cross Abstract: Federated learning (FL) offers a promising approach to privacy-preserving clinical risk prediction, but its deployment remains limited by restricted data sharing, client heterogeneity, class imbalance, and the lack of realistic tabular electronic health record (EHR) benchmarks.

By Akarsh K Nair, Muhammad Arifur Rahman, Nicholas Shopland, Andy Burton, Jun He, Yuan Shen, David Baldwin, Emma O'Dowd, Amna Burzic, Mufti Mahmud, David J. Brown
arXiv Machine Learning
1d ago

Rethinking Anonymity Claims in Synthetic Data Generation: A Model-Centric Privacy Attack Perspective

The paper argues that evaluating anonymity in synthetic data generation must focus on the generative model rather than just the resulting dataset. It interprets GDPR definitions of personal data and anonymization under realistic model-access scenarios, mapping these to state‑of‑the‑art privacy attacks. The authors conclude that synthetic data alone is insufficient for anonymization, and that Differential Privacy offers stronger protection than Similarity‑based Privacy Metrics.

By Georgi Ganev, Emiliano De Cristofaro
arXiv AI
3d ago

Aegis: Generative Gradient Masking for Privacy-Preserving Medical Federated Learning

Aegis is a client‑side defense for medical federated learning that protects against model inversion attacks by adding a masking gradient derived from locally synthesized data. The method exploits the fact that attacks fail when the effective batch size exceeds the model’s leakage capacity, turning this bottleneck into a privacy guarantee. Experiments on MNIST, CIFAR‑10, and MedMNIST datasets show that Aegis neutralizes state‑of‑the‑art attacks while preserving model accuracy and adding only modest overhead.

By Chaoyu Zhang, Shanghao Shi, Heng Jin, Ning Wang, Y. Thomas Hou, Wenjing Lou
arXiv AI
Jul 23

Recovering Clinical Utility Under Differential Privacy: Empirical Validation of Adaptive Federated Aggregation on Heterogeneous Cardiovascular Datasets

arXiv:2607. 19403v1 Announce Type: cross Abstract: Validating federated learning frameworks on real clinical data is an essential step between proof-of-concept demonstrations in controlled synthetic environments and deployment in real multicenter healthcare settings.

By Rodrigo Tertulino, Laercio Alencar, Ricardo Almeida
arXiv Machine Learning
Aug 28

Private and interpretable clinical prediction with quantum-inspired tensor train models

The paper demonstrates that publicly available clinical machine learning models, such as logistic regression (LR), pose significant privacy risks because attackers can recover model parameters and identify training cohorts through various membership inference attacks. The authors show that even small cohorts can be reliably identified and that common practices like cross-validation can worsen the risk. To mitigate this, they propose a quantum-inspired defense that tensorizes discretized models into tensor trains (TTs), which obfuscates parameters, preserves accuracy, and maintains interpretability while providing black‑box protection comparable to Differential Privacy.

By Jos\'e Ram\'on Pareja Monturiol, Juliette Sinnott, Roger G. Melko, Mohammad Kohandel