arXiv Machine Learning

EigenLI: Spectral Approximations to Late Interaction

EigenLI introduces a spectral approximation framework that compresses late‑interaction representations by identifying document‑specific low‑dimensional subspaces. By selecting dominant eigendirections, it constructs reduced interaction representations that outperform clustering‑based pooling methods on ColBERTv2 and AnswerAI‑ColBERT‑small. The framework also yields EigenLI‑SV, a single‑vector ANN‑compatible representation that consistently surpasses comparable surrogates such as MUVERA across multiple datasets and text models.

arXiv AI
5d ago

A Manifold-Aware Topic Modeling Approach via Rank-Based Prototypes

The paper introduces MARETopic, a training‑free framework that identifies topics by selecting rank‑based prototype documents from pretrained embeddings. By projecting embeddings onto a low‑dimensional manifold and building ranked neighborhood lists, a greedy algorithm picks exactly K exemplar texts whose neighborhoods cover the corpus. Two variants—MARETopic_Corr, which uses a query‑performance predictor and rank correlation, and MARETopic_Diff, which employs a rank‑based diffusion matrix—achieve higher purity and NMI on benchmark datasets and run significantly faster, while also improving topic coherence and vocabulary diversity through a novel Maximal Marginal Relevance step.

By Thiago C\'esar Castilho Almeida, Daniel Carlos Guimar\~aes Pedronette
arXiv AI
Sep 18

Randomized SVD Approximations for Spectral Co-Clustering of Word-Document Matrices

The paper introduces two randomized approaches to accelerate spectral co‑clustering of word‑document matrices: one based on randomized SVD via random projection, and another combining partial SVD with element‑wise random sampling. Experiments on real and synthetic data show both methods cut runtime compared to full SVD, with the projection technique offering more consistent performance across varied sparsity levels, while the sampling method excels on denser matrices. The study highlights that the choice of approximation should align with the data’s structural properties.

By Fateme Mazdarani, Carlos Toxtli
arXiv Machine Learning
Jul 28

SpecFormer: Mitigating Embedding and Attention Collapse via Spectral-Aware Transformer for Recommendation

arXiv:2607. 24025v1 Announce Type: cross Abstract: Transformer architectures have achieved remarkable success across diverse domains; however, directly applying their standard self-attention mechanism to recommendation often yields suboptimal performance, sometimes even trailing behind well-designed simple recommendation models.

By Yu Cui, Yi Xu, Jiahao Wang, Hao Zhang, Yu Zhang, Xiaoyi Zeng, Can Wang, Jinxin Hu, Jiawei Chen
arXiv Machine Learning
Sep 1

Effective Graph and Rank-based Contextual Embeddings for Textual and Multimedia Data

arXiv:2608.29001v1 Announce Type: new Abstract: In a data-driven world, efficiently organizing and mapping relationships between objects is crucial. Graphs are powerful tools for modeling these conne...

By Thiago C\'esar Castilho Almeida, Gustavo Rosseto Let\'icio, Lucas Pascotti Valem, Andr\'e Freitas, Daniel Carlos Guimar\~aes Pedronette
arXiv Computation and Language
Sep 23

KaLM-Reranker-V1: Fast but Not Late Interaction for Compressed Document Reranking

arXiv:2606.22807v3 Announce Type: replace Abstract: As retrieval systems scale, effective and efficient reranking becomes increasingly important. However, most existing encoder- and decoder-based rer...

By Xinping Zhao, Jiaxin Xu, Ziqi Dai, Xin Zhang, Huiyao Chen, Shouzheng Huang, Xianhao Xiong, Danyu Tang, Xinshuo Hu, Guohong Fu, Meishan Zhang, Baotian Hu