arXiv Machine Learning By Jihwan Kim, Chenglin Fan

DP-Muon: Differentially Private Optimization via Matrix-Orthogonalized Momentum

Read the original on arXiv Machine Learning →

The paper introduces DP-Muon, a differentially private optimization method that incorporates matrix‑orthogonalized momentum. It employs standard global per‑example clipping and releases a single Gaussian‑noised gradient per step, with matrix and auxiliary updates treated as post‑processing. The authors analyze the mean distortion introduced when fresh Gaussian noise passes through a nonlinear matrix map, deriving exact Gaussian heat identities and showing that for a smooth Newton‑Schulz map, the conditional output bias is reduced from second to fourth order in the noise scale. They also establish matrix‑block stationarity bounds, quantify orthogonalization error, and provide criteria for improving the upper bound, while a separate inequality captures the impact of auxiliary Adam updates. Experiments on GPT‑2 at various privacy targets demonstrate that DP‑Muon configurations outperform Adam baselines in test negative log‑likelihood.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.