SuperPCA is a new algorithm for high‑dimensional principal component analysis that exploits an approximate eigenspace of the sample covariance matrix. The authors show that the subspace spanned by several leading eigenvectors contains useful signal information long before individual eigenvectors converge, and they derive posteriori bounds on the angle between this subspace and the true signal subspace. By using only a small number of subsampled coordinates, SuperPCA can achieve up to a ten‑fold improvement in accuracy over classical PCA while reducing data acquisition costs, especially when the signals are approximately sparse.
By Irina-Beatrice Haas, Maike Meier, Yuji Nakatsukasa, Taejun Park
arXiv:2606. 03553v1 Announce Type: cross Abstract: While principal component analysis (PCA) is a fundamental tool for dimensionality reduction, its dense representations make it ill-suited for high-dimensional data.
By David V\"avinggren, Francis Bach, Andr\'e M. H. Teixeira, Dave Zachariah, Ant\^onio H. Ribeiro
arXiv:2607. 23198v1 Announce Type: new Abstract: We propose Variance-Preserving Orthogonal Selection (VPOS), a greedy framework for unsupervised feature selection that operates in the weighted PCA loading space.
By Baran Koseoglu, Berrin Yanikoglu
arXiv:2608.29362v1 Announce Type: cross
Abstract: Sparse computations are fundamental to scientific computing, graph analytics, and machine learning, yet their performance is highly sensitive to the...
By Ruifeng Zhang, Xipeng Shen
arXiv:2601. 10199v2 Announce Type: replace Abstract: Multivariate data often exhibit complex dependencies that violate the assumption of isotropic residual noise.
By Antonio Briola, Marwin Schmidt, Fabio Caccioli, Carlos Ros Perez, James Singleton, Christian Michler, Tomaso Aste
arXiv:2609.14815v1 Announce Type: cross
Abstract: This paper introduces a novel framework for Regularized Multivariate Functional Principal Component Analysis (ReMFPCA) via Functional Singular Value...
By Yue Zhao, Hossein Haghbin, Rebecca Sanders, Mehdi Maadooliat
arXiv:2606. 06233v1 Announce Type: cross Abstract: Principal component analysis (PCA) is one of the most widely used unsupervised dimension reduction techniques.
By Benedikt Seiter, Anya Fries, Julius von K\"ugelgen, Jonas Peters
arXiv:2306. 14851v5 Announce Type: replace-cross Abstract: Given a high-dimensional covariate matrix and a response vector, ridge-regularized sparse linear regression selects a subset of features that explains the relationship between covariates and the response in an interpretable manner.
By Ryan Cory-Wright, Andr\'es G\'omez
The paper studies streaming principal component analysis under a robust setting where the covariance matrix can vary within a temporal uncertainty set, rather than being fixed. It establishes fundamental convergence limits for any algorithm that recovers principal components and analyzes the noisy power method and Oja's algorithm, showing that the noisy power method achieves rate‑optimal convergence in this setting. Numerical experiments on synthetic and real‑world data confirm the theoretical findings.
By Daniel Bienstock, Minchan Jeong, Apurv Shukla, Se-Young Yun
Calendar-Structured Sparse Principal Component Analysis (Calendar-SPCA) is a new method that learns low-dimensional representations of long-term electricity consumption data by explicitly incorporating daily, weekly, and annual calendar cycles. It uses an L1 penalty and graph total variation to produce sparse, locally coherent, and directly interpretable latent factors. In experiments on the GoiEner and Low Carbon London smart‑meter datasets, Calendar-SPCA retains most of the variance of standard PCA while achieving high sparsity and clear calendar‑aligned structures.
By Carlos Quesada-Granja, Tony Castillo-Calzadilla, Carlos Rizo-Maestre
The paper introduces robust multi-task procedures for principal component analysis that leverage similarity across tasks to enhance eigenspace estimation while remaining resilient to outlier tasks. It establishes non-asymptotic convergence rates and demonstrates that the methods achieve minimax optimal performance across various regimes. One procedure, based on matrix-depth, attains optimal error dependence on the proportion of outlier tasks, addressing a key challenge in robust multi-task learning.
By Dali Liu, Haolei Weng
arXiv:2602. 08913v3 Announce Type: replace Abstract: In underdetermined regression and classification problems, multiple feature subsets often yield equivalent predictive performance.
By Kate\v{r}ina Henclov\'a, V\'aclav \v{S}m\'idl