Spectral graph clustering with inhomogeneous latent geometry
arXiv:2608. 11321v1 Announce Type: cross Abstract: We study spectral clustering in the presence of a confounding latent geometry.
arXiv:2608. 11321v1 Announce Type: cross Abstract: We study spectral clustering in the presence of a confounding latent geometry.
The paper investigates Partial Least Squares (PLS) in high-dimensional settings, focusing on a model where two data matrices share a low-rank latent structure plus individual-specific components. By analyzing the singular vectors of the cross‑covariance matrix with random matrix theory, the authors derive asymptotic characterizations of how well the estimated latent directions align with the true ones. They show that the PLS variant based on Singular Value Decomposition (PLS‑SVD) outperforms separate principal component analysis in detecting the common latent subspace, while also identifying regimes where PLS‑SVD behaves counter‑intuitively or reaches fundamental limits.
arXiv:2512.22282v2 Announce Type: replace-cross Abstract: Across fields such as machine learning, social science, and geology, considerable attention has been given to models that factorize a nonnega...
arXiv:2607. 18883v1 Announce Type: cross Abstract: A central aim of unsupervised learning is to uncover latent factors that explain dependencies among observations.
EigenLI introduces a spectral approximation framework that compresses late‑interaction representations by identifying document‑specific low‑dimensional subspaces. By selecting dominant eigendirections, it constructs reduced interaction representations that outperform clustering‑based pooling methods on ColBERTv2 and AnswerAI‑ColBERT‑small. The framework also yields EigenLI‑SV, a single‑vector ANN‑compatible representation that consistently surpasses comparable surrogates such as MUVERA across multiple datasets and text models.
arXiv:2608. 05243v1 Announce Type: cross Abstract: Factorized generative models commonly regularize a latent style variable z_s by matching its marginal distribution to a fixed Gaussian prior and interpret this as evidence that the style representation is independent of class information.
arXiv:2608. 08704v1 Announce Type: cross Abstract: Kernel spectral clustering with a single bandwidth can be inadequate for data exhibiting multiple characteristic pairwise-distance scales, a problem particularly prevalent in the high-dimensional regime.
arXiv:2608.28799v1 Announce Type: cross Abstract: Separable nonnegative matrix factorization (SNMF) has been widely used for low-rank representation and clustering of nonnegative data, owing to its a...
arXiv:2601. 18128v2 Announce Type: replace-cross Abstract: High-dimensional data often exhibit variation that can be captured by lower-dimensional factors.
The paper introduces a new dimensionality reduction technique that enhances nearest‑neighbour relationships to estimate high‑information projections. It constructs a matrix encoding local covariance via nearest‑neighbour pairs and shows that, under standard regularity conditions, this matrix consistently estimates the Density Information Matrix (DIM), a non‑parametric analogue of the Fisher Information Matrix. The authors also demonstrate the method’s practical usefulness for clustering and outlier detection.
arXiv:2605. 16836v2 Announce Type: replace-cross Abstract: Hypergraphs provide a principled framework for modeling polyadic interactions, with applications in recommendation systems, social networks, and molecular modeling.
arXiv:2606. 28854v1 Announce Type: cross Abstract: The common factor analytic model is related to Helmholtz and Boltzmann machines, can be conceived as a linear autoencoder, or can be thought of as a single-hidden-layer generative neural network.