The paper compares two popular data‑integration techniques—Stack‑SVD, which concatenates datasets before performing singular value decomposition, and SVD‑Stack, which first decomposes each dataset separately and then aggregates the leading singular vectors. By deriving exact asymptotic performance expressions and phase transitions in a proportional regime, the authors show that neither method uniformly dominates the other when unweighted, but optimally weighted Stack‑SVD outperforms optimally weighted SVD‑Stack when the low‑rank signal is fully shared. They also demonstrate that SVD‑Stack can excel with partially shared components and provide practical algorithms for estimating optimal weights, supported by simulations and genomic experiments.
By Tavor Z. Baharav, Phillip B. Nicol, Rafael A. Irizarry, Rong Ma
arXiv:2609.14815v1 Announce Type: cross
Abstract: This paper introduces a novel framework for Regularized Multivariate Functional Principal Component Analysis (ReMFPCA) via Functional Singular Value...
By Yue Zhao, Hossein Haghbin, Rebecca Sanders, Mehdi Maadooliat
arXiv:2605.15240v2 Announce Type: replace-cross
Abstract: This paper investigates the critical role of eigenalignments between the kernel matrix and learning targets in achieving robust generalizatio...
By Yang Liu, Ernest Fokoue, Richard Lange, Daniel Krutz
The paper introduces robust multi-task procedures for principal component analysis that leverage similarity across tasks to enhance eigenspace estimation while remaining resilient to outlier tasks. It establishes non-asymptotic convergence rates and demonstrates that the methods achieve minimax optimal performance across various regimes. One procedure, based on matrix-depth, attains optimal error dependence on the proportion of outlier tasks, addressing a key challenge in robust multi-task learning.
By Dali Liu, Haolei Weng
SuperPCA is a new algorithm for high‑dimensional principal component analysis that exploits an approximate eigenspace of the sample covariance matrix. The authors show that the subspace spanned by several leading eigenvectors contains useful signal information long before individual eigenvectors converge, and they derive posteriori bounds on the angle between this subspace and the true signal subspace. By using only a small number of subsampled coordinates, SuperPCA can achieve up to a ten‑fold improvement in accuracy over classical PCA while reducing data acquisition costs, especially when the signals are approximately sparse.
By Irina-Beatrice Haas, Maike Meier, Yuji Nakatsukasa, Taejun Park
The paper introduces a new dimensionality reduction technique that enhances nearest‑neighbour relationships to estimate high‑information projections. It constructs a matrix encoding local covariance via nearest‑neighbour pairs and shows that, under standard regularity conditions, this matrix consistently estimates the Density Information Matrix (DIM), a non‑parametric analogue of the Fisher Information Matrix. The authors also demonstrate the method’s practical usefulness for clustering and outlier detection.
By David P. Hofmeyr