arXiv Machine Learning By Beatrix M. G. Nielsen, Andreas Grivas

What Cosine Similarity of Label Representations Can and Cannot Tell us

Read the original on arXiv Machine Learning →

arXiv:2603. 29488v2 Announce Type: replace Abstract: Cosine similarity is often used to measure the similarity of vector representations of neural network models.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 9

Similarity-Distance-Magnitude Activations

arXiv:2509. 12760v5 Announce Type: replace Abstract: We introduce the Similarity-Distance-Magnitude (SDM) activation function, a more robust and interpretable formulation of the standard softmax activation function, adding Similarity (i.

By Allen Schmaltz
arXiv Machine Learning
Sep 24

Binary Classification from Coupled Pairwise Labels

The paper introduces SD-Pcomp learning, a binary classification framework that jointly utilizes Similarity/Dissimilarity (SD) labels and Pairwise Comparison (Pcomp) labels from instance pairs. It proposes an objective function that can be decomposed into either an SD estimator plus ordering information or a Pcomp estimator plus pair-type information, thereby integrating complementary relational cues. Experiments on eight datasets demonstrate that combining both label types improves classification accuracy and AUC compared to using either alone or a simple convex combination.

By Tomoya Tate, Kosuke Sugiyama, Masato Uchida
arXiv Machine Learning
Sep 10

Learning Kernels by Alignment for Multiclass Bayes Classification

The paper introduces a framework that learns the kernel used in kernel methods through alignment, leveraging the Collaborative Learning and Inference (CLaI) approach. It demonstrates that CLaI can be interpreted as a kernel alignment process and that its inference stage is equivalent to kernel Bayes classification with Parzen-window density estimation. By replacing cosine similarity with a learned Mahalanobis distance, the authors extend CLaI to multiclass classification, achieving higher accuracy, faster convergence, and lower calibration error on datasets such as CIFAR-10, PathMNIST, and SleepEDF, while also showing connections to Gaussian processes and competitive calibration in sepsis prediction.

By Hollan Haule, Alfredo Gonzalez-Sulser, Javier Escudero