Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direction is equally meaningful, but there is no reason...
arXiv:2609.27988v1 Announce Type: cross
Abstract: Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direc...
By Andrew Bond, Ege Erdem \"Ozl\"u, Tuna \c{C}imen, Ilkin Umut Melanlioglu, Tolga Birdal, Erkut Erdem, Aykut Erdem
arXiv:2601. 11618v2 Announce Type: replace-cross Abstract: Geometric Attention (GA) specifies an attention layer by four independent inputs: a finite carrier (what indices are addressable), an evidence-kernel rule (how masked proto-scores and a link induce nonnegative weights), a probe family (which observables are treated as admissible), and an anchor/update rule (which representative kernel is selected and how it is applied).
By Luis Rosario Freytes
arXiv:2609.15083v1 Announce Type: new
Abstract: Mixed-curvature representation learning seeks to capture rich geometric structures that cannot be adequately modeled by a single curvature regime. Exis...
By Xingrun Li, Yusuke Mukuta, Xin Yang, Yinyu Ye, Tatsuya Harada
arXiv:2608.19021v2 Announce Type: replace
Abstract: Global Covariance Pooling (GCP) improves deep networks by capturing second-order feature statistics, and is especially effective for fine-grained r...
By Md Rifat Ur Rahman, Md Raihan Khan, Md Sakib Hossain Shovon, Pietro Li\`o, Mohammad Ali Moni
The paper introduces Mahalanobis-Based Multi-Head Attention for Complex State Propagation (MHA‑CSP), a new attention mechanism that replaces the standard dot‑product with a Mahalanobis distance‑based RBF kernel. This approach enables infinite‑dimensional feature space attention without extra parameters, allows direct construction of Tree Attention via LogSumExp correction, and incorporates an attention meshing mechanism for cross‑head collaboration. Experiments show that with only 119K parameters and teacher forcing applied only at the final hidden state, MHA‑CSP outperforms Transformer and GCN baselines on long‑sequence state tracking tasks, demonstrating efficient structured reasoning.
By Xiaohe Li