arXiv:2508.04800v2 Announce Type: replace-cross
Abstract: We introduce a novel privatization framework for high-dimensional controlled variable selection. Our framework enables rigorous False Discove...
By Yuxuan Tao, Adel Javanmard
The paper addresses the challenge of selecting external datasets for private transfer learning by modeling high‑dimensional regression with heterogeneous sources and a weighted ridge estimator. It relies solely on aggregated statistics and offers privacy guarantees under $
ho$‑zero‑concentrated differential privacy for labels or both features and labels. A deterministic equivalent of test error is derived, enabling optimization of hyperparameters and decision‑making about the utility of private external data without accessing individual records.
By Filip Kova\v{c}evi\'c, Edwige Cyffers, Stefano Sarao Mannelli, Marco Mondelli
arXiv:2602. 17284v2 Announce Type: replace Abstract: We consider the privacy amplification properties of a sampling scheme in which a user's data isused in $k$ steps chosen randomly and uniformly from a sequence (or set) of $t$ steps.
By Vitaly Feldman, Moshe Shenfeld
arXiv:2609.37344v1 Announce Type: cross
Abstract: Data reconstruction attacks have empirically been successful in recovering training samples from learned models, raising privacy concerns and motivat...
By Max Cairney-Leeming, Simone Bombari, Marco Mondelli
arXiv:2503. 10945v3 Announce Type: replace-cross Abstract: Current practices for reporting differential privacy (DP) guarantees for machine learning (ML) algorithms such as DP-SGD provide an incomplete and potentially misleading picture.
By Juan Felipe Gomez, Bogdan Kulynych, Georgios Kaissis, Flavio P. Calmon, Jamie Hayes, Borja Balle, Antti Honkela
The paper presents a new analysis of Oja's algorithm for streaming principal component analysis (PCA) that works without any eigengap assumptions, achieving near‑optimal rates and matching lower bounds. It extends the results to a Rayleigh quotient notion of approximate PCA, resolving an open question, and applies the findings to provide gap‑free differentially private PCA guarantees for sub‑Gaussian data. The analysis relies solely on a second‑moment bound of stochastic updates, avoiding the almost‑sure bounds used in previous work.
By Anming Gu, Syamantak Kumar, Kevin Tian, Chutong Yang
arXiv:2608.28934v1 Announce Type: new
Abstract: Differential privacy (DP) has traditionally been used to provide theoretical upper bounds on an algorithm's stability to changing its training data. In...
By Saloni Modi, Srivi Balaji, Yusong Zhu, Gautam Kamath, Kevin Tian
arXiv:2605. 11170v2 Announce Type: replace Abstract: Noise-based certified machine unlearning currently faces a hard ceiling: the noise magnitude required to certify unlearning typically destroys model utility, particularly for large-scale deletion requests.
By Ahmed Mehdi Inane, Vincent Quirion, Gintare Karolina Dziugaite, Ioannis Mitliagkas
arXiv:2602. 01607v3 Announce Type: replace-cross Abstract: Differentially private synthetic data enables the sharing and analysis of sensitive datasets while providing rigorous privacy guarantees for individual contributors.
By Rundong Ding, Yiyun He, Yizhe Zhu
arXiv:2601. 10237v3 Announce Type: replace Abstract: Differentially Private Stochastic Gradient Descent (DP-SGD) is the dominant paradigm for private training, but its fundamental limitations under worst-case adversarial privacy definitions remain poorly understood.
By Murat Bilgehan Ertan, Marten van Dijk
arXiv:2303. 07152v3 Announce Type: replace-cross Abstract: Achieving optimal statistical performance while ensuring the privacy of personal data is a challenging yet crucial objective in modern data analysis.
By T. Tony Cai, Yichen Wang, Linjun Zhang
arXiv:2606. 09582v1 Announce Type: new Abstract: Recent work argues for using Gaussian differential privacy (GDP) to report the privacy guarantees in privacy-preserving machine learning.
By Bogdan Kulynych, Antti Honkela