The paper tackles two key gaps in streaming PCA using Oja's algorithm: it establishes sharp operator‑norm convergence for general‑rank subspaces under sub‑Gaussian data, and it provides distributional inference for the resulting subspace estimator. The authors remove non‑vanishing remainder terms from existing analyses, achieving rates that match minimax bounds in both dense‑tail and sparse‑tail regimes. They further develop a linearization of Oja’s iterates, enabling high‑dimensional Gaussian approximations and an online multiplier bootstrap for practical inference.
By Haoshu Xu, Hongzhe Li
SuperPCA is a new algorithm for high‑dimensional principal component analysis that exploits an approximate eigenspace of the sample covariance matrix. The authors show that the subspace spanned by several leading eigenvectors contains useful signal information long before individual eigenvectors converge, and they derive posteriori bounds on the angle between this subspace and the true signal subspace. By using only a small number of subsampled coordinates, SuperPCA can achieve up to a ten‑fold improvement in accuracy over classical PCA while reducing data acquisition costs, especially when the signals are approximately sparse.
By Irina-Beatrice Haas, Maike Meier, Yuji Nakatsukasa, Taejun Park
arXiv:2602. 02190v2 Announce Type: replace-cross Abstract: A common approach to perform PCA on probability measures is to embed them into a Hilbert space where standard functional PCA techniques apply.
By Gachon Erell, J\'er\'emie Bigot, Elsa Cazelles
arXiv:2609.05796v1 Announce Type: cross
Abstract: Principal component analysis (PCA) can rotate away from its population target when a covariance matrix is estimated from limited data. We introduce d...
By Qiang Sun
The paper studies streaming principal component analysis under a robust setting where the covariance matrix can vary within a temporal uncertainty set, rather than being fixed. It establishes fundamental convergence limits for any algorithm that recovers principal components and analyzes the noisy power method and Oja's algorithm, showing that the noisy power method achieves rate‑optimal convergence in this setting. Numerical experiments on synthetic and real‑world data confirm the theoretical findings.
By Daniel Bienstock, Minchan Jeong, Apurv Shukla, Se-Young Yun
arXiv:2606. 14533v1 Announce Type: new Abstract: Principal Component Analysis (PCA) preserves variance, not the information needed to detect rare catastrophic events.
By Hamidou Tembine
arXiv:2505. 19925v2 Announce Type: replace-cross Abstract: The sample covariance matrix is a cornerstone of multivariate statistics, but it is highly sensitive to outliers.
By Fabio Centofanti, Mia Hubert, Peter J. Rousseeuw
arXiv:2609. 36142v1 Announce Type: cross Abstract: In Bayesian inference problems with non-Gaussian observation noise, the posterior is only as accurate as the noise density, and gradient-based samplers need that density and its gradient evaluable pointwise, whether from an explicit expression or from code, and without an inner solve.
By Joshua Chen, Peter Jan van Leeuwen
arXiv:2604. 03146v2 Announce Type: replace-cross Abstract: We study high-dimensional convex empirical risk minimization (ERM) under general non-Gaussian data designs.
By Chiheb Yaakoubi, Cosme Louart, Malik Tiomoko, Zhenyu Liao
arXiv:2511. 11927v2 Announce Type: replace-cross Abstract: Principal Component Analysis (PCA) is a standard tool for extracting a low-rank signal from noisy observations.
By Urte Adomaityte, Gabriele Sicuro, Pierpaolo Vivo
arXiv:2607. 21823v1 Announce Type: new Abstract: We show that, up to isotropic scaling, the Gaussian RBF reproducing kernel Hilbert space (RKHS) is asymptotically isometric to Euclidean space in the large bandwidth limit.
By Sergio A. Alvarez
arXiv:2607. 16638v1 Announce Type: cross Abstract: Principal component regression (PCR) regularizes high-dimensional prediction by choosing a spectral cutoff, but rank selection cannot correct systematic inflation of the retained empirical eigenvalues.
By Peng Zhao