The paper investigates the use of the squared norm of a whitened foundation‑model embedding as a training‑free likelihood surrogate. It shows that the apparent Gaussianity of whitened coordinates stems from the projection central limit theorem, not from a true joint Gaussian distribution, and that the norm is systematically over‑dispersed compared to a Gaussian reference. The authors explain that whitening reverses the encoder’s spectral hierarchy, concentrating norm contributions in near‑degenerate directions dominated by noise, and propose interpreting the squared norm as a Mahalanobis measure of semantic atypicality rather than a log‑likelihood.
By Mohammed Ahnouch, Lotfi Elaachak
The standard way to compare two text embeddings is cosine similarity. Scattered studies report that a different metric does better, but never pin down the geometric condition that decides when, or why.
arXiv:2504. 16318v3 Announce Type: replace Abstract: Cosine similarity is a standard comparison rule for learned representations in information retrieval, natural language processing, computer vision, and multimodal learning.
By Kisung You
arXiv:2607. 19393v1 Announce Type: cross Abstract: While auditing a perturbation-based OOD detector on a document benchmark, we recorded an AUROC of 0.
By Vishnu Bindu Balachandran
arXiv:2602. 19393v2 Announce Type: replace Abstract: Steck, Ekanadham, and Kallus [arXiv:2403.
By Taha Bouhsine
Latent visual reasoning (LVR) inserts supervised latent tokens between perception and answer generation in vision-language models (VLMs). The field uses alignment between these latents and their visual targets, i.
The paper introduces a framework that learns the kernel used in kernel methods through alignment, leveraging the Collaborative Learning and Inference (CLaI) approach. It demonstrates that CLaI can be interpreted as a kernel alignment process and that its inference stage is equivalent to kernel Bayes classification with Parzen-window density estimation. By replacing cosine similarity with a learned Mahalanobis distance, the authors extend CLaI to multiclass classification, achieving higher accuracy, faster convergence, and lower calibration error on datasets such as CIFAR-10, PathMNIST, and SleepEDF, while also showing connections to Gaussian processes and competitive calibration in sepsis prediction.
By Hollan Haule, Alfredo Gonzalez-Sulser, Javier Escudero
arXiv:2609.36307v1 Announce Type: new
Abstract: Partial Least Squares (PLS) regression extracts a few outcome-aligned directions in a high-dimensional X and is widely used across applied science, but...
By Pawe{\l} Lenartowicz, Hubert Plisiecki
The paper introduces PLSP (Pre-hoc Liminal Space Profiling), an anticipatory framework for predicting out-of-distribution (OOD) data before inference. It proposes a dataset‑independent metric called the CREDibility Score (CREDS) and introduces credibility curves and heat maps to analyze a model’s maximum credibility and behavior across datasets. Experiments on multiple datasets show that CREDS can improve model robustness to OOD prediction.
By Vipul Bansal, Himanshu Buckchash, Balasubramanian Raman, Deepak Dhungana
arXiv:2503. 05169v2 Announce Type: replace Abstract: Applying machine learning to increasingly high-dimensional problems with sparse or biased training data increases the risk that a model is used on inputs outside its training domain.
By Felix Krumbiegel, Juniper Tyree, Michael Boy, Petri Clusius, Andreas Rupp
arXiv:2609.22522v1 Announce Type: new
Abstract: Cosine similarity is widely used to analyze transformer representations, implicitly assuming that similarity reflects task-relevant structure. We study...
By Yu Sun, Mengyin Lu, Cong Feng, Guangming Lu, Huimin Han
The paper extends mechanistic interpretability of large language models by modeling concepts as low‑dimensional non‑linear manifolds rather than linear subspaces. It introduces a concept‑based alignment (CBA) score to compare these manifolds across layers and models, revealing block structures in intermediate layers, a shift from syntax‑dominated to mixed syntactic‑semantic concepts, and training‑dependent multilingual sharing. The study also shows that alignment patterns differ across model families and training stages, with adjacent stages aligning more closely than distant ones.
By Tido Specht, Elias Benedict Krey, Nils Neukirch, Nils Strodthoff