arXiv:2606. 10333v1 Announce Type: new Abstract: Credit risk prediction is a critical problem in the consumer credit industry.
By Hongzhe Zhang, Jiarong Xu, Jing He, Xiao Fang
QuanText is a training‑free, large‑language‑model‑agnostic mechanism for releasing textual datasets that protects dataset‑level secrets such as the proportion of records with a particular diagnosis or gender. It perturbs both the secret distribution and correlated attribute distributions by selecting candidate release distributions close to the private empirical distribution and rewriting each text sample to match the chosen distribution using attribute‑related snippets. The method is inspired by the Statistic Maximal Leakage framework and, under idealized conditions, satisfies an SML guarantee, while empirical evaluations show a superior privacy‑utility trade‑off compared to existing data generation baselines.
By Shuaiqi Wang, Zinan Lin, Giulia Fanti
The paper explores how missing data can inherently enhance privacy in machine learning. By integrating missingness into a differential privacy framework, the authors demonstrate that the absence of certain features can amplify privacy guarantees without altering the underlying algorithm. This reveals a previously overlooked interaction between data incompleteness and formal privacy protections.
By Simon Roburin (LPSM), Rafa{\"e}l Pinot (LPSM), Erwan Scornet (LPSM)
arXiv:2503. 23536v3 Announce Type: replace-cross Abstract: Unlearnable data (ULD) has emerged as an innovative defense technique to prevent machine learning models from learning meaningful patterns from specific data, thus protecting data privacy and security.
By Jiahao Li, Yiqiang Chen, Yunbing Xing, Yang Gu, Xiangyuan Lan
arXiv:2608. 13773v1 Announce Type: cross Abstract: Neural networks are increasingly deployed in high-stakes applications with growing privacy leakage concerns.
By Leonardo Magliolo, Vito Paolo Pastore, Giuseppe Valenzise, Enzo Tartaglione
arXiv:2307. 13127v3 Announce Type: replace-cross Abstract: Data used to train predictive models via empirical risk minimization (ERM) often contain sensitive personal information.
By Spencer Giddens, Yiwang Zhou, Kevin R. Krull, Tara M. Brinkman, Peter X. K. Song, Fang Liu