arXiv Machine Learning By Victor L\'eger, Florent Chatelain

High-Dimensional Partial Least Squares: Spectral Analysis and Fundamental Limitations

Read the original on arXiv Machine Learning →

The paper investigates Partial Least Squares (PLS) in high-dimensional settings, focusing on a model where two data matrices share a low-rank latent structure plus individual-specific components. By analyzing the singular vectors of the cross‑covariance matrix with random matrix theory, the authors derive asymptotic characterizations of how well the estimated latent directions align with the true ones. They show that the PLS variant based on Singular Value Decomposition (PLS‑SVD) outperforms separate principal component analysis in detecting the common latent subspace, while also identifying regimes where PLS‑SVD behaves counter‑intuitively or reaches fundamental limits.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
6d ago

Stacked SVD or SVD stacked? A Random Matrix Theory perspective on data integration

The paper compares two popular data‑integration techniques—Stack‑SVD, which concatenates datasets before performing singular value decomposition, and SVD‑Stack, which first decomposes each dataset separately and then aggregates the leading singular vectors. By deriving exact asymptotic performance expressions and phase transitions in a proportional regime, the authors show that neither method uniformly dominates the other when unweighted, but optimally weighted Stack‑SVD outperforms optimally weighted SVD‑Stack when the low‑rank signal is fully shared. They also demonstrate that SVD‑Stack can excel with partially shared components and provide practical algorithms for estimating optimal weights, supported by simulations and genomic experiments.

By Tavor Z. Baharav, Phillip B. Nicol, Rafael A. Irizarry, Rong Ma
arXiv Statistics ML
Sep 18

Robust Multi-Task Learning for Principal Component Analysis

The paper introduces robust multi-task procedures for principal component analysis that leverage similarity across tasks to enhance eigenspace estimation while remaining resilient to outlier tasks. It establishes non-asymptotic convergence rates and demonstrates that the methods achieve minimax optimal performance across various regimes. One procedure, based on matrix-depth, attains optimal error dependence on the proportion of outlier tasks, addressing a key challenge in robust multi-task learning.

By Dali Liu, Haolei Weng
arXiv Machine Learning
Sep 23

SuperPCA: subspace analysis and an efficient algorithm for high-dimensional PCA

SuperPCA is a new algorithm for high‑dimensional principal component analysis that exploits an approximate eigenspace of the sample covariance matrix. The authors show that the subspace spanned by several leading eigenvectors contains useful signal information long before individual eigenvectors converge, and they derive posteriori bounds on the angle between this subspace and the true signal subspace. By using only a small number of subsampled coordinates, SuperPCA can achieve up to a ten‑fold improvement in accuracy over classical PCA while reducing data acquisition costs, especially when the signals are approximately sparse.

By Irina-Beatrice Haas, Maike Meier, Yuji Nakatsukasa, Taejun Park
arXiv Machine Learning
Aug 27

Efficient Estimation of High Information Projections using Nearest Neighbours

The paper introduces a new dimensionality reduction technique that enhances nearest‑neighbour relationships to estimate high‑information projections. It constructs a matrix encoding local covariance via nearest‑neighbour pairs and shows that, under standard regularity conditions, this matrix consistently estimates the Density Information Matrix (DIM), a non‑parametric analogue of the Fisher Information Matrix. The authors also demonstrate the method’s practical usefulness for clustering and outlier detection.

By David P. Hofmeyr