arXiv:2603. 11308v3 Announce Type: replace Abstract: Principal Component Analysis (PCA) is a cornerstone of dimensionality reduction, yet its classical formulation relies critically on second-order moments and is therefore fragile in the presence of heavy-tailed data and impulsive noise.
By Mario Sayde, Christopher Khater, Jihad Fahs, Ibrahim Abou-Faycal
SuperPCA is a new algorithm for high‑dimensional principal component analysis that exploits an approximate eigenspace of the sample covariance matrix. The authors show that the subspace spanned by several leading eigenvectors contains useful signal information long before individual eigenvectors converge, and they derive posteriori bounds on the angle between this subspace and the true signal subspace. By using only a small number of subsampled coordinates, SuperPCA can achieve up to a ten‑fold improvement in accuracy over classical PCA while reducing data acquisition costs, especially when the signals are approximately sparse.
By Irina-Beatrice Haas, Maike Meier, Yuji Nakatsukasa, Taejun Park
arXiv:2607. 16638v1 Announce Type: cross Abstract: Principal component regression (PCR) regularizes high-dimensional prediction by choosing a spectral cutoff, but rank selection cannot correct systematic inflation of the retained empirical eigenvalues.
By Peng Zhao
arXiv:2609.05796v1 Announce Type: cross
Abstract: Principal component analysis (PCA) can rotate away from its population target when a covariance matrix is estimated from limited data. We introduce d...
By Qiang Sun
arXiv:2606. 06233v1 Announce Type: cross Abstract: Principal component analysis (PCA) is one of the most widely used unsupervised dimension reduction techniques.
By Benedikt Seiter, Anya Fries, Julius von K\"ugelgen, Jonas Peters
arXiv:2505. 19925v2 Announce Type: replace-cross Abstract: The sample covariance matrix is a cornerstone of multivariate statistics, but it is highly sensitive to outliers.
By Fabio Centofanti, Mia Hubert, Peter J. Rousseeuw
arXiv:2607. 23682v1 Announce Type: new Abstract: Early warning of extreme market volatility is central to financial risk management, but actionable events are rare, nonstationary, and often triggered by exogenous information shocks.
By Jin Qian, Zhangzhi Xiong, Mingrui Li, Zhen Liu
arXiv:2607. 11947v1 Announce Type: cross Abstract: Typical semi-supervised learning (SSL) methods rely on distributional assumptions, and their performance degrades when these are violated.
By Yushi Hirose, Hiroo Irobe, Takafumi Kanamori
arXiv:2603. 28257v2 Announce Type: replace-cross Abstract: KAN-PCA is an autoencoder that uses a KAN as encoder and a linear map as decoder.
By David Breazu
arXiv:2602. 10680v2 Announce Type: replace-cross Abstract: Many real-world datasets contain hidden structure that cannot be detected by simple linear correlations between input features.
By Vicente Conde Mendes, Lorenzo Bardone, C\'edric Koller, Jorge Medina Moreira, Vittorio Erba, Emanuele Troiani, Lenka Zdeborov\'a
The paper investigates Partial Least Squares (PLS) in high-dimensional settings, focusing on a model where two data matrices share a low-rank latent structure plus individual-specific components. By analyzing the singular vectors of the cross‑covariance matrix with random matrix theory, the authors derive asymptotic characterizations of how well the estimated latent directions align with the true ones. They show that the PLS variant based on Singular Value Decomposition (PLS‑SVD) outperforms separate principal component analysis in detecting the common latent subspace, while also identifying regimes where PLS‑SVD behaves counter‑intuitively or reaches fundamental limits.
By Victor L\'eger, Florent Chatelain
arXiv:2606. 25007v1 Announce Type: new Abstract: Financial fraud detection in digital banking requires reasoning over multiple heterogeneous event streams -- transactions, login sessions, risk signals -- that individually appear benign but collectively reveal fraudulent patterns.
By Mohammadamin Dashti Moghaddam, Nick Sciarrilli