arXiv Machine Learning

Robust Streaming PCA

The paper studies streaming principal component analysis under a robust setting where the covariance matrix can vary within a temporal uncertainty set, rather than being fixed. It establishes fundamental convergence limits for any algorithm that recovers principal components and analyzes the noisy power method and Oja's algorithm, showing that the noisy power method achieves rate‑optimal convergence in this setting. Numerical experiments on synthetic and real‑world data confirm the theoretical findings.

arXiv Machine Learning
Aug 20

Inference and Uncertainty Quantification for Streaming $r$-PCA

The paper tackles two key gaps in streaming PCA using Oja's algorithm: it establishes sharp operator‑norm convergence for general‑rank subspaces under sub‑Gaussian data, and it provides distributional inference for the resulting subspace estimator. The authors remove non‑vanishing remainder terms from existing analyses, achieving rates that match minimax bounds in both dense‑tail and sparse‑tail regimes. They further develop a linearization of Oja’s iterates, enabling high‑dimensional Gaussian approximations and an online multiplier bootstrap for practical inference.

By Haoshu Xu, Hongzhe Li
arXiv Machine Learning
Sep 23

SuperPCA: subspace analysis and an efficient algorithm for high-dimensional PCA

SuperPCA is a new algorithm for high‑dimensional principal component analysis that exploits an approximate eigenspace of the sample covariance matrix. The authors show that the subspace spanned by several leading eigenvectors contains useful signal information long before individual eigenvectors converge, and they derive posteriori bounds on the angle between this subspace and the true signal subspace. By using only a small number of subsampled coordinates, SuperPCA can achieve up to a ten‑fold improvement in accuracy over classical PCA while reducing data acquisition costs, especially when the signals are approximately sparse.

By Irina-Beatrice Haas, Maike Meier, Yuji Nakatsukasa, Taejun Park
arXiv Machine Learning
Jul 27

Heavy-Tailed Principal Component Analysis

arXiv:2603. 11308v3 Announce Type: replace Abstract: Principal Component Analysis (PCA) is a cornerstone of dimensionality reduction, yet its classical formulation relies critically on second-order moments and is therefore fragile in the presence of heavy-tailed data and impulsive noise.

By Mario Sayde, Christopher Khater, Jihad Fahs, Ibrahim Abou-Faycal
arXiv Machine Learning
Sep 23

Gap-Free Streaming PCA Beyond Rank-One Updates: Near-Optimal Rates and Applications to Differential Privacy

The paper presents a new analysis of Oja's algorithm for streaming principal component analysis (PCA) that works without any eigengap assumptions, achieving near‑optimal rates and matching lower bounds. It extends the results to a Rayleigh quotient notion of approximate PCA, resolving an open question, and applies the findings to provide gap‑free differentially private PCA guarantees for sub‑Gaussian data. The analysis relies solely on a second‑moment bound of stochastic updates, avoiding the almost‑sure bounds used in previous work.

By Anming Gu, Syamantak Kumar, Kevin Tian, Chutong Yang
arXiv Machine Learning
Jun 5

Anchor PCA

arXiv:2606. 06233v1 Announce Type: cross Abstract: Principal component analysis (PCA) is one of the most widely used unsupervised dimension reduction techniques.

By Benedikt Seiter, Anya Fries, Julius von K\"ugelgen, Jonas Peters