The paper studies operator learning on function spaces using encoder–decoder architectures. It shows that as input and output resolutions grow, the induced kernels converge to a limiting kernel, enabling regularity assumptions independent of resolution. The authors derive upper and lower bounds for regularized stochastic gradient descent, extend the analysis to neural networks via the limiting neural tangent kernel, and provide error bounds and complexity guarantees for various kernel and encoding constructions.
By Lei Shi, Jia-Qi Yang, Ding-Xuan Zhou
arXiv:2609.27988v1 Announce Type: cross
Abstract: Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direc...
By Andrew Bond, Ege Erdem \"Ozl\"u, Tuna \c{C}imen, Ilkin Umut Melanlioglu, Tolga Birdal, Erkut Erdem, Aykut Erdem
arXiv:2606. 03879v1 Announce Type: cross Abstract: As foundation models scale toward fusing more heterogeneous visual streams, understanding how diverse encoders interact under joint training becomes a prerequisite for principled design.
By Wei Ding, Yudong Zhang, Ruobing Xie, Xingwu Sun, Jiansheng Chen, Yu Wang
The paper investigates the use of the squared norm of a whitened foundation‑model embedding as a training‑free likelihood surrogate. It shows that the apparent Gaussianity of whitened coordinates stems from the projection central limit theorem, not from a true joint Gaussian distribution, and that the norm is systematically over‑dispersed compared to a Gaussian reference. The authors explain that whitening reverses the encoder’s spectral hierarchy, concentrating norm contributions in near‑degenerate directions dominated by noise, and propose interpreting the squared norm as a Mahalanobis measure of semantic atypicality rather than a log‑likelihood.
By Mohammed Ahnouch, Lotfi Elaachak
Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direction is equally meaningful, but there is no reason...
The paper investigates how input encodings constrain the set of contrasts a predictor can reproduce, even when no individual contrast is forced to zero. By computing the attainable contrast space from an encoder’s equivalence classes and a fixed contrast design—without using labels, loss, or a fitted model—the authors derive an empirical error floor for any unrestricted decoder on those classes. Experiments on a 140‑rectangle siRNA interaction panel show that a graph neural network’s training‑only feature mask reduces the rank of interaction contrasts from 140 to 72, creating a floor of 0.009980 (14.6% of the fitted model’s interaction squared error). Removing the mask eliminates the floor but only marginally improves MSE, while restoring chemistry columns recovers full rank. A separate RNA‑splicing predictor with an injective encoding achieves full rank and a zero floor, illustrating that the encoding itself, not the model, limits recoverable contrast space.
By Zahra Khodagholi, Niloofar Yousefi