arXiv:2508. 10017v2 Announce Type: replace-cross Abstract: Federated Learning (FL) presents a groundbreaking approach for collaborative health research, allowing model training on decentralized data while safeguarding patient privacy.
By Rodrigo Tertulino
arXiv:2606. 06990v1 Announce Type: new Abstract: The generation of high-fidelity synthetic Electronic Health Records (EHR) is crucial for advancing medical research while preserving patient privacy.
By Jalen Jiang, Chufan Gao, Ethan Rasmussen, Stephen Z. Xie, Jimeng Sun
arXiv:2607. 19524v1 Announce Type: cross Abstract: Federated learning (FL) offers a promising approach to privacy-preserving clinical risk prediction, but its deployment remains limited by restricted data sharing, client heterogeneity, class imbalance, and the lack of realistic tabular electronic health record (EHR) benchmarks.
By Akarsh K Nair, Muhammad Arifur Rahman, Nicholas Shopland, Andy Burton, Jun He, Yuan Shen, David Baldwin, Emma O'Dowd, Amna Burzic, Mufti Mahmud, David J. Brown
arXiv:2606. 16952v2 Announce Type: replace-cross Abstract: The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets.
By Kareem Amin, Rudrajit Das, Alessandro Epasto, Adel Javanmard, Dennis Kraft, M\'onica Ribero, Sergei Vassilvitskii
The paper argues that evaluating anonymity in synthetic data generation must focus on the generative model rather than just the resulting dataset. It interprets GDPR definitions of personal data and anonymization under realistic model-access scenarios, mapping these to state‑of‑the‑art privacy attacks. The authors conclude that synthetic data alone is insufficient for anonymization, and that Differential Privacy offers stronger protection than Similarity‑based Privacy Metrics.
By Georgi Ganev, Emiliano De Cristofaro
arXiv:2606. 08259v1 Announce Type: new Abstract: This paper investigates the problem of generating synthetic tabular data with differential privacy (DP) guarantees, enabling data sharing in sensitive domains.
By Toan Tran, Arturs Backurs, Zinan Lin, Victor Reis, Li Xiong, Sergey Yekhanin
arXiv:2602. 05833v2 Announce Type: replace Abstract: There is a need for synthetic training and test datasets that replicate statistical distributions of original datasets without compromising their confidentiality.
By Laura Plein, Alexi Turcotte, Arina Hallemans, Andreas Zeller
arXiv:2606. 16952v1 Announce Type: cross Abstract: The rapid adoption of generative AI and Large Language Models (LLMs) has spurred interest in synthetic data as a privacy-preserving alternative to sensitive real-world datasets.
By Kareem Amin, Rudrajit Das, Alessandro Epasto, Adel Javanmard, Dennis Kraft, M\'onica Ribero, Sergei Vassilvitskii
Aegis is a client‑side defense for medical federated learning that protects against model inversion attacks by adding a masking gradient derived from locally synthesized data. The method exploits the fact that attacks fail when the effective batch size exceeds the model’s leakage capacity, turning this bottleneck into a privacy guarantee. Experiments on MNIST, CIFAR‑10, and MedMNIST datasets show that Aegis neutralizes state‑of‑the‑art attacks while preserving model accuracy and adding only modest overhead.
By Chaoyu Zhang, Shanghao Shi, Heng Jin, Ning Wang, Y. Thomas Hou, Wenjing Lou
arXiv:2607. 19403v1 Announce Type: cross Abstract: Validating federated learning frameworks on real clinical data is an essential step between proof-of-concept demonstrations in controlled synthetic environments and deployment in real multicenter healthcare settings.
By Rodrigo Tertulino, Laercio Alencar, Ricardo Almeida
The paper demonstrates that publicly available clinical machine learning models, such as logistic regression (LR), pose significant privacy risks because attackers can recover model parameters and identify training cohorts through various membership inference attacks. The authors show that even small cohorts can be reliably identified and that common practices like cross-validation can worsen the risk. To mitigate this, they propose a quantum-inspired defense that tensorizes discretized models into tensor trains (TTs), which obfuscates parameters, preserves accuracy, and maintains interpretability while providing black‑box protection comparable to Differential Privacy.
By Jos\'e Ram\'on Pareja Monturiol, Juliette Sinnott, Roger G. Melko, Mohammad Kohandel
arXiv:2608.28934v1 Announce Type: new
Abstract: Differential privacy (DP) has traditionally been used to provide theoretical upper bounds on an algorithm's stability to changing its training data. In...
By Saloni Modi, Srivi Balaji, Yusong Zhu, Gautam Kamath, Kevin Tian