arXiv Machine Learning

High-Dimensional Asymptotics and Dataset Selection for Private Transfer Learning

The paper addresses the challenge of selecting external datasets for private transfer learning by modeling high‑dimensional regression with heterogeneous sources and a weighted ridge estimator. It relies solely on aggregated statistics and offers privacy guarantees under $ ho$‑zero‑concentrated differential privacy for labels or both features and labels. A deterministic equivalent of test error is derived, enabling optimization of hyperparameters and decision‑making about the utility of private external data without accessing individual records.

arXiv Machine Learning
Oct 2

Certification-Based Differentially Private Learning

The paper extends the abstract gradient training (AGT) framework to provide tighter differential privacy guarantees for both private prediction and private learning. It introduces Abstract Gradient Sampling (AGS) to analyze privacy in continuous, unbounded regression and offers theoretical and empirical evidence that these methods yield tighter bounds than global-sensitivity baselines, even in previously unbounded settings. The authors also demonstrate that their private learning algorithm can outperform standard private learners under comparable conditions.

By Mihnea Ghitu, Matthew Wicker
arXiv Machine Learning
Sep 11

Label Differential Privacy via Aggregation

arXiv:2310. 10092v4 Announce Type: replace Abstract: This paper explores the use of linear aggregation to protect the privacy of sensitive training labels through the concept of \emph{label differential privacy} (label-DP) while maintaining regression task utility.

By Anand Brahmbhatt, Rishi Saket, Shreyas Havaldar, Anshul Nasery, Yukti Makhija, Aravindan Raghuveer
Hugging Face Trending Papers
Jun 1

IntraShuffler: A Privacy Preserving Framework for Heterogeneous DP Federated Learning

Heterogeneous Differential Privacy (HDP) in Federated Learning (FL) allows clients to select individual privacy budgets ($\varepsilon_i$) according to institutional policies and data sensitivity. In practice, many HDP-FL systems employ $\varepsilon$-aware server aggregation to improve model utility by re-weighting client updates according to their declared privacy budgets.