arXiv Machine Learning

Closed-Form Noise Calibration Against Membership Inference for Random-Allocation DP-SGD

arXiv:2610. 09651v1 Announce Type: new Abstract: DP-SGD protects training data by adding Gaussian noise to clipped gradients.

arXiv Machine Learning
4d ago

Trade-off Functions for DP-SGD with Subsampling based on Random Allocation: Tight Upper and Lower Bounds

arXiv:2605. 06259v3 Announce Type: replace Abstract: Within the $f$-DP framework, we derive a tight analysis of the trade-off function for Differentially Private Stochastic Gradient Descent (DP-SGD) with subsampling based on random allocation in which each sample is independently assigned to exactly one of $M$ minibatches per epoch, each minibatch corresponding to one of the $M$ SGD rounds within a single epoch.

By Marten van Dijk, Murat Bilgehan Ertan
arXiv Machine Learning
Sep 14

Membership Inference via Pairwise Likelihood Ratios

The paper introduces Pairwise Likelihood MIA (PL‑MIA), a unified membership inference attack that combines a Gaussian likelihood‑ratio statistic with population calibration and the Cauchy combination test. PL‑MIA generates p‑values from pairwise comparisons between a query point and reference points, then aggregates these continuous signals using the Cauchy test to preserve evidence strength. Experiments show that PL‑MIA surpasses strong baselines, boosting true positive rates by over 25% in low‑false‑positive settings, thereby validating the theoretical advantages of the proposed statistical framework.

By Shengjie Niu, Zebin Yun, Yeheng Ge, Jian Huang
arXiv AI
Sep 10

PAC-Private Autoregressive Generation: Calibrating Noise to Ensemble Disagreement

The paper introduces PAC‑Private Autoregressive Generation, a method that calibrates noise based on ensemble disagreement across overlapping ‘worlds’ of a private corpus, thereby extending PAC privacy from classification to text generation. By training adapters on a frozen public model and using posterior‑weighted disagreement to add noise only when predictions vary, the approach achieves strong privacy guarantees while preserving most of the fine‑tuning benefit. Experiments on WikiText‑103 with GPT‑2‑small show 74 % of the fine‑tuning gain retained with a per‑token budget of 2⁻³², and membership‑inference success bounded to 51.08 % after one million tokens, outperforming PMixED under matched conditions.

By Mina Mirzadehsarcheshmeh, Amir Keyvan Khandani