arXiv:2605. 06259v3 Announce Type: replace Abstract: Within the $f$-DP framework, we derive a tight analysis of the trade-off function for Differentially Private Stochastic Gradient Descent (DP-SGD) with subsampling based on random allocation in which each sample is independently assigned to exactly one of $M$ minibatches per epoch, each minibatch corresponding to one of the $M$ SGD rounds within a single epoch.
By Marten van Dijk, Murat Bilgehan Ertan
arXiv:2601. 10237v3 Announce Type: replace Abstract: Differentially Private Stochastic Gradient Descent (DP-SGD) is the dominant paradigm for private training, but its fundamental limitations under worst-case adversarial privacy definitions remain poorly understood.
By Murat Bilgehan Ertan, Marten van Dijk
The paper introduces Pairwise Likelihood MIA (PL‑MIA), a unified membership inference attack that combines a Gaussian likelihood‑ratio statistic with population calibration and the Cauchy combination test. PL‑MIA generates p‑values from pairwise comparisons between a query point and reference points, then aggregates these continuous signals using the Cauchy test to preserve evidence strength. Experiments show that PL‑MIA surpasses strong baselines, boosting true positive rates by over 25% in low‑false‑positive settings, thereby validating the theoretical advantages of the proposed statistical framework.
By Shengjie Niu, Zebin Yun, Yeheng Ge, Jian Huang
arXiv:2609. 22783v1 Announce Type: new Abstract: We study differentially private covariance estimation in operator norm for mean-zero sub-Gaussian distributions with unknown covariance support and at most $k$ nonzero entries per row.
By Zihan Zhang
arXiv:2607. 19510v1 Announce Type: new Abstract: Modern LLM deployments use a number of implementation choices and inference optimizations (e.
By Eric Price, Kevin Tian, Zhiyang Xun, Yusong Zhu
arXiv:2606. 01527v2 Announce Type: replace Abstract: Machine unlearning is motivated by legal and user-facing requirements to remove the influence of individuals' data from trained models, such as the right to be forgotten.
By Matthew Regehr, Gautam Kamath, Andrew Lowy
The paper introduces PAC‑Private Autoregressive Generation, a method that calibrates noise based on ensemble disagreement across overlapping ‘worlds’ of a private corpus, thereby extending PAC privacy from classification to text generation. By training adapters on a frozen public model and using posterior‑weighted disagreement to add noise only when predictions vary, the approach achieves strong privacy guarantees while preserving most of the fine‑tuning benefit. Experiments on WikiText‑103 with GPT‑2‑small show 74 % of the fine‑tuning gain retained with a per‑token budget of 2⁻³², and membership‑inference success bounded to 51.08 % after one million tokens, outperforming PMixED under matched conditions.
By Mina Mirzadehsarcheshmeh, Amir Keyvan Khandani
arXiv:2601. 14033v2 Announce Type: replace Abstract: Machine learning models are increasingly served behind APIs.
By Xiaochen Zhu, Mayuri Sridhar, Srinivas Devadas
arXiv:2509. 25003v3 Announce Type: replace Abstract: Membership inference attacks (MIAs) against Diffusion Models (DMs) raise pressing privacy concerns by revealing whether a sample was part of the training set.
By Mingxing Rao, Bowen Qu, Daniel Moyer
arXiv:2605. 11170v2 Announce Type: replace Abstract: Noise-based certified machine unlearning currently faces a hard ceiling: the noise magnitude required to certify unlearning typically destroys model utility, particularly for large-scale deletion requests.
By Ahmed Mehdi Inane, Vincent Quirion, Gintare Karolina Dziugaite, Ioannis Mitliagkas
arXiv:2608.27782v1 Announce Type: cross
Abstract: Memorization in large language models is measured through a zoo of definitions whose formal relations are unknown, and differential privacy (DP) is t...
By Xujun Che, Depeng Xu, Shuhan Yuan
arXiv:2603. 11799v2 Announce Type: replace Abstract: Membership inference attacks (MIAs) are becoming standard tools for auditing the privacy of machine learning models.
By Rickard Br\"annvall