arXiv Machine Learning

Privacy-Preserving Credit Risk Prediction with Alternative Data

arXiv:2606. 10333v1 Announce Type: new Abstract: Credit risk prediction is a critical problem in the consumer credit industry.

arXiv Machine Learning
1d ago

Rethinking Anonymity Claims in Synthetic Data Generation: A Model-Centric Privacy Attack Perspective

The paper argues that evaluating anonymity in synthetic data generation must focus on the generative model rather than just the resulting dataset. It interprets GDPR definitions of personal data and anonymization under realistic model-access scenarios, mapping these to state‑of‑the‑art privacy attacks. The authors conclude that synthetic data alone is insufficient for anonymization, and that Differential Privacy offers stronger protection than Similarity‑based Privacy Metrics.

By Georgi Ganev, Emiliano De Cristofaro
arXiv Machine Learning
Sep 14

A Differentially Private Federated Proximal Optimization Framework for Customer Churn Prediction in Heterogeneous Federated Telecom Networks

The paper introduces a Differentially Private Federated Proximal (DP‑FedProx) framework for predicting customer churn in telecom networks. It compares DP‑FedProx with FedAvg, DP‑FedAvg, FedProx, and several centralized/local models on two public churn datasets, showing that DP‑FedProx consistently outperforms FedAvg variants and matches the best centralized model with only a slight accuracy loss while offering privacy guarantees. SHAP analysis reveals that DP‑FedProx prioritizes revenue‑group features, highlighting its practical balance between performance and privacy.

By Joydeb Kumar Sana, Subrata Chakraborty, M M Manjurul Islam
arXiv Machine Learning
Jun 30

Efficient Unlearning with Privacy Guarantees

arXiv:2507. 04771v2 Announce Type: replace-cross Abstract: Privacy protection laws, such as the GDPR, grant individuals the right to request the forgetting of their personal data not only from databases but also from machine learning (ML) models trained on them.

By Josep Domingo-Ferrer, Najeeb Jebreel, David S\'anchez
Hugging Face Trending Papers
Jun 3

Federated Learning for Multi-Center Sepsis Early Prediction with Privacy-Preserving

Privacy-sensitive and distributed characteristics of multi-center medical data bring severe obstacles to centralized modeling for accurate early prediction of sepsis. Federated learning (FL) has attracted growing attention as a promising framework for collaborative model development, as it allows multiple institutions to jointly train predictive models without directly sharing or centralizing raw data.

arXiv Machine Learning
Sep 17

QuanText: Protecting Dataset-Level Secrets in Textual Data Sharing

QuanText is a training‑free, large‑language‑model‑agnostic mechanism for releasing textual datasets that protects dataset‑level secrets such as the proportion of records with a particular diagnosis or gender. It perturbs both the secret distribution and correlated attribute distributions by selecting candidate release distributions close to the private empirical distribution and rewriting each text sample to match the chosen distribution using attribute‑related snippets. The method is inspired by the Statistic Maximal Leakage framework and, under idealized conditions, satisfies an SML guarantee, while empirical evaluations show a superior privacy‑utility trade‑off compared to existing data generation baselines.

By Shuaiqi Wang, Zinan Lin, Giulia Fanti