arXiv Machine Learning

Gap-free Differentially Private PCA for Gaussian Data

The paper presents a gap‑free differentially private algorithm for performing principal component analysis on Gaussian data. It addresses the PCA problem while ensuring privacy guarantees without relying on a spectral gap assumption. The work is announced on arXiv with the identifier 2609.31614v1.

arXiv Machine Learning
Sep 23

Gap-Free Streaming PCA Beyond Rank-One Updates: Near-Optimal Rates and Applications to Differential Privacy

The paper presents a new analysis of Oja's algorithm for streaming principal component analysis (PCA) that works without any eigengap assumptions, achieving near‑optimal rates and matching lower bounds. It extends the results to a Rayleigh quotient notion of approximate PCA, resolving an open question, and applies the findings to provide gap‑free differentially private PCA guarantees for sub‑Gaussian data. The analysis relies solely on a second‑moment bound of stochastic updates, avoiding the almost‑sure bounds used in previous work.

By Anming Gu, Syamantak Kumar, Kevin Tian, Chutong Yang
arXiv Machine Learning
Sep 11

DP-Muon: Differentially Private Optimization via Matrix-Orthogonalized Momentum

The paper introduces DP-Muon, a differentially private optimization method that incorporates matrix‑orthogonalized momentum. It employs standard global per‑example clipping and releases a single Gaussian‑noised gradient per step, with matrix and auxiliary updates treated as post‑processing. The authors analyze the mean distortion introduced when fresh Gaussian noise passes through a nonlinear matrix map, deriving exact Gaussian heat identities and showing that for a smooth Newton‑Schulz map, the conditional output bias is reduced from second to fourth order in the noise scale. They also establish matrix‑block stationarity bounds, quantify orthogonalization error, and provide criteria for improving the upper bound, while a separate inequality captures the impact of auxiliary Adam updates. Experiments on GPT‑2 at various privacy targets demonstrate that DP‑Muon configurations outperform Adam baselines in test negative log‑likelihood.

By Jihwan Kim, Chenglin Fan