arXiv:2602. 23785v2 Announce Type: replace Abstract: We investigate the identifiability of nonlinear canonical correlation analysis (CCA) in a multi-view setup, in which each view is generated by applying an unknown nonlinear map to a linear mixture of shared latent variables plus view-private noise.
By Zhiwei Han, Stefan Matthes, Hao Shen
arXiv:2505.18918v4 Announce Type: replace-cross
Abstract: Principal component analysis (PCA) is a key tool in the field of data dimensionality reduction. Various methods have been proposed to extend...
By Javier Salazar Cavazos, Jeffrey A Fessler, Laura Balzano
arXiv:2411. 12438v2 Announce Type: replace-cross Abstract: We develop a new approach for clustering non-spherical (i.
By Prashanti Anderson, Mitali Bafna, Rares-Darius Buhai, Pravesh K. Kothari, David Steurer
arXiv:2608.29102v1 Announce Type: new
Abstract: This paper develops a unified theoretical framework showing that a broad family of clustering methods, including k-means, fuzzy c-means, kernel k-means...
By Angshul Majumdar
The paper introduces a new dimensionality reduction technique that enhances nearest‑neighbour relationships to estimate high‑information projections. It constructs a matrix encoding local covariance via nearest‑neighbour pairs and shows that, under standard regularity conditions, this matrix consistently estimates the Density Information Matrix (DIM), a non‑parametric analogue of the Fisher Information Matrix. The authors also demonstrate the method’s practical usefulness for clustering and outlier detection.
By David P. Hofmeyr
arXiv:2511. 07109v2 Announce Type: replace-cross Abstract: Nonnegative matrix factorization (NMF) is a linear dimensionality reduction technique for nonnegative data, with applications such as hyperspectral unmixing and topic modeling.
By Junjun Pan, Valentin Leplat, Michael Ng, Nicolas Gillis
arXiv:2212. 07944v4 Announce Type: replace Abstract: We study a multi-factor block model for variable clustering and connect it to regularized subspace clustering through a distributionally robust version of nodewise regression.
By Kaizheng Wang, Xiao Xu, Xun Yu Zhou
arXiv:2608. 14215v1 Announce Type: new Abstract: Constrained optimization extends classical optimization by integrating side information, making it widely applicable across scientific and engineering domains.
By Johanna Hillebrand, Jan H\"ockendorff, J\"urgen Kusche, Kelin Luo, Heiko R\"oglin, Melanie Schmidt, Christian Sohler, Bernd Uebbing
arXiv:2608. 08704v1 Announce Type: cross Abstract: Kernel spectral clustering with a single bandwidth can be inadequate for data exhibiting multiple characteristic pairwise-distance scales, a problem particularly prevalent in the high-dimensional regime.
By Zeqin Lin, Guangming Pan, Zhixiang Zhang, Yinbing Zhou
arXiv:2607. 21039v1 Announce Type: new Abstract: Spectral methods are among the most widely used techniques for community detection, clustering, and graph learning.
By Zhuan Liang, Zheng Zhai
The paper investigates Partial Least Squares (PLS) in high-dimensional settings, focusing on a model where two data matrices share a low-rank latent structure plus individual-specific components. By analyzing the singular vectors of the cross‑covariance matrix with random matrix theory, the authors derive asymptotic characterizations of how well the estimated latent directions align with the true ones. They show that the PLS variant based on Singular Value Decomposition (PLS‑SVD) outperforms separate principal component analysis in detecting the common latent subspace, while also identifying regimes where PLS‑SVD behaves counter‑intuitively or reaches fundamental limits.
By Victor L\'eger, Florent Chatelain
The paper compares two popular data‑integration techniques—Stack‑SVD, which concatenates datasets before performing singular value decomposition, and SVD‑Stack, which first decomposes each dataset separately and then aggregates the leading singular vectors. By deriving exact asymptotic performance expressions and phase transitions in a proportional regime, the authors show that neither method uniformly dominates the other when unweighted, but optimally weighted Stack‑SVD outperforms optimally weighted SVD‑Stack when the low‑rank signal is fully shared. They also demonstrate that SVD‑Stack can excel with partially shared components and provide practical algorithms for estimating optimal weights, supported by simulations and genomic experiments.
By Tavor Z. Baharav, Phillip B. Nicol, Rafael A. Irizarry, Rong Ma