arXiv AI

Low-Interaction-Rank Learning: Unifying Multiplicative Dual-Encoder Heads

arXiv:2608. 11661v1 Announce Type: cross Abstract: A multiplicative dual-encoder network computes a real-valued output for a pair of inputs as the inner product of their separate encodings.

arXiv Machine Learning
Sep 15

Resolution-Independent Analysis of Encoder--Decoder Operator Learning via Limiting Kernels

The paper studies operator learning on function spaces using encoder–decoder architectures. It shows that as input and output resolutions grow, the induced kernels converge to a limiting kernel, enabling regularity assumptions independent of resolution. The authors derive upper and lower bounds for regularized stochastic gradient descent, extend the analysis to neural networks via the limiting neural tangent kernel, and provide error bounds and complexity guarantees for various kernel and encoding constructions.

By Lei Shi, Jia-Qi Yang, Ding-Xuan Zhou
arXiv Machine Learning
Sep 22

Whitening Inverts the Hierarchy: What the Norm of a Whitened Embedding Measures

The paper investigates the use of the squared norm of a whitened foundation‑model embedding as a training‑free likelihood surrogate. It shows that the apparent Gaussianity of whitened coordinates stems from the projection central limit theorem, not from a true joint Gaussian distribution, and that the norm is systematically over‑dispersed compared to a Gaussian reference. The authors explain that whitening reverses the encoder’s spectral hierarchy, concentrating norm contributions in near‑degenerate directions dominated by noise, and propose interpreting the squared norm as a Mahalanobis measure of semantic atypicality rather than a log‑likelihood.

By Mohammed Ahnouch, Lotfi Elaachak
arXiv AI
4d ago

Representable but Unlearned: Encoding Rank and the Interaction-Prediction Floor

The paper investigates how input encodings constrain the set of contrasts a predictor can reproduce, even when no individual contrast is forced to zero. By computing the attainable contrast space from an encoder’s equivalence classes and a fixed contrast design—without using labels, loss, or a fitted model—the authors derive an empirical error floor for any unrestricted decoder on those classes. Experiments on a 140‑rectangle siRNA interaction panel show that a graph neural network’s training‑only feature mask reduces the rank of interaction contrasts from 140 to 72, creating a floor of 0.009980 (14.6% of the fitted model’s interaction squared error). Removing the mask eliminates the floor but only marginally improves MSE, while restoring chemistry columns recovers full rank. A separate RNA‑splicing predictor with an injective encoding achieves full rank and a zero floor, illustrating that the encoding itself, not the model, limits recoverable contrast space.

By Zahra Khodagholi, Niloofar Yousefi
arXiv Computer Vision
2d ago

Platonic Task Arithmetic

arXiv:2610.00929v1 Announce Type: cross Abstract: Models specialized for the same task converge to similar behavior, yet the parameter updates that produce it share no common coordinate system, so we...

By Junghwan Park, Woojin Cho
arXiv AI
Sep 10

MoEMB: Scaling Universal Multimodal Embeddings with Efficient Mixture-of-Experts Models

MoEMB introduces a mixture‑of‑experts (MoE) approach to scale universal multimodal embeddings (UME) without increasing the size of the output vector or relying on autoregressive decoding. By expanding encoder capacity along the expert axis, MoEMB achieves state‑of‑the‑art performance on MMEB‑V2 and MRMR benchmarks with only 3 B active parameters, outperforming TTE‑based methods that use more than four times as many active parameters and require significantly more compute. The paper also presents the first comprehensive study of adaptive computation for MoE‑based embeddings, exploring training‑time and inference‑time strategies to further improve efficiency for large‑scale retrieval and recommendation systems.

By Xuanming Cui, Shlok Kumar Mishra, Wentao Bao, Aashu Singh, Zihao Wang, Xiangjun Fan, Jun Xiao, Ser-Nam Lim, Jianpeng Cheng